A field playbook · Managed service providers

Scale Service Delivery with Governed AI

A field playbook for managed service providers to serve more clients per engineer without letting quality or security slip.

First edition · 2026 · jimiige.com
Jimi Ige · AI Governance, Deployment and Adoption
Why this, and how to read it

Most MSPs grow the same way: sign the client, stretch the engineers, and hope the desk absorbs the load until the next hire lands. It works until the scarce thing runs out, and the scarce thing is never technicians or tools. It is senior engineering attention, and the runbook knowledge that lives in a few heads and resigns when they do. The economics sharpen the problem: the fee per seat is fixed at signature, so every saved engineer-hour is margin, every recurring ticket is a leak, and every escalation that reaches a senior engineer spends the most expensive attention the firm owns. Meanwhile AI has already arrived on the desk, in suggested replies and summarized tickets, whether the firm decided anything about it or not.

This playbook is about operations, not technology. The product is an MSP that serves more seats per engineer, keeps resolution knowledge as firm property instead of personal property, and walks into every QBR with evidence. AI is the enabling technology. Governance is the reason the result can be trusted inside client environments, and the reason the improvement survives its first difficult quarter. One rule anchors everything here: an accountable engineer stays on every consequential action, and AI never holds standing permission to change a client environment. It is written for the person who owns the outcome: the operator who has to grow the client base without the service, or the security posture, slipping.

A note on method. Every number in this playbook is cited to a named primary source and carries its own caveat, and where the honest evidence is a gap, the gap is stated instead of filled. For MSPs the largest gap is stated up front: no approved primary source in this library covers MSP margin, utilization, or seat economics, so those are taught qualitatively. No invented clients, no vendor arithmetic, no borrowed payback periods. The pattern is specific enough to test against your own desk, with a 90-day way to run that test.

Capacity
Serve more seats per engineer, by moving resolution knowledge and recurring work onto governed, reusable rails.
Outcome 01
Quality
Raise desk throughput while escalation tiers, change approval, and accountable engineers stay exactly where they are.
Outcome 02
Trust
Credentials, client boundaries, and security posture protected by design, and provable when the questionnaire arrives.
Outcome 03

Three commitments, no hockey sticks. Each chapter ends with where judgment beats the tool, because the client is not buying tickets closed. They are buying the confidence to stop thinking about IT.

Novice tilt
across the causal studies of AI assistance in knowledge work, the gains have been reported concentrating among less experienced workers on supported tasks, which is a tier-one leverage story and a supervision one at the same time
Noy & Zhang; Dell'Acqua et al.; Brynjolfsson, Li & Raymond · a cross-study inference, never an MSP measurement
Scaffolds
two external structures exist to hang MSP AI controls on: NIST's AI Risk Management Framework with its Generative AI Profile, and ISO/IEC 42001, the first international AI management system standard
NIST AI RMF 1.0 and NIST.AI.600-1; ISO/IEC 42001:2023 · voluntary, not mandates and not a liability safe harbor
88%
of organizations report regular AI use in at least one function, while only 39 percent report enterprise-level EBIT impact, in a self-reported global survey of 1,993 respondents
McKinsey State of AI, 2025

Frame these precisely. The experience-gap pattern is a cross-study inference drawn from task-level experiments in other kinds of work, not a measurement of an MSP desk, and the frameworks are voluntary scaffolds: the existence of a standard is not evidence that anything works. The adoption figures are self-reported and skew to large enterprises. And the number this page cannot give you is the industry one: no approved primary source in this library covers MSP margin, utilization, or seat economics, so this playbook teaches seat economics qualitatively and asks you to baseline your own desk. Full source notes close the playbook.

Chapter 1

The per-seat machine: how an MSP actually makes money

The principle

Strip the branding and an MSP is a subscription on engineering attention. The client pays a fixed monthly fee per seat or endpoint; the firm promises outcomes: things patched, tickets resolved, users unblocked, environments secure. Revenue is decided at signature. Cost is decided every day afterward, one ticket at a time. That inversion is the whole economics: every saved engineer-hour is margin, and every recurring ticket is a leak that runs for the life of the agreement. Three dials govern the model: how many seats each engineer can support, how much ticket load each seat generates, and which tier of engineer touches each ticket. All three bend around the same two constraints. Senior engineers, the ones who can untangle the hard escalation and design the fix that stops a ticket from coming back, do not scale by hiring. And the knowledge that keeps the desk fast, the runbooks, the environment quirks, the fix that worked last time, lives mostly in heads, and it resigns when they do.

The trap

The trap is buying tools while leaving the operating model alone. The stack grows: another monitoring agent, another automation module, another assistant bolted to the ticketing system. A year later the desk types faster, and the same tickets still recur, the same escalations still reach the same two seniors, and onboarding a new client still takes the best engineer offline for a month. The licenses were real. Seats per engineer never moved, because nothing about how the desk resolves, documents, escalates, or learns was redesigned.

Every saved engineer-hour is margin. Every recurring ticket is a leak.

The checklist

  • Name the constraint on your desk: intake noise, tier-one resolution, escalation load, or documentation debt. They are different problems with different fixes.
  • Follow one week of senior-engineer hours and mark which of them only that engineer could have spent.
  • Count the tickets the desk has already solved before: same issue, same client, or same fix rediscovered from scratch.
  • Price each agreement against the effort it actually consumes, not the effort the quote assumed.
  • Choose workflows to improve, not tools to buy. A tool dropped on an unchanged workflow returns almost nothing.

Where judgment beats the tool

An effort analysis shows where engineer time goes. It cannot say which recurring ticket is an engineering problem, which is a client-behavior problem, and which is a pricing problem wearing a technical costume. Deciding whether to fix the environment, coach the client, or reprice the agreement is an owner's call, and each answer sends the same ticket somewhere different.

Chapter 2

Four outcomes that matter, and one that does not

The principle

Four outcomes justify this whole program. Grow the client base without proportional engineering hires, so revenue can move faster than payroll. Increase capacity per engineer without letting quality or security slip, because in this business one bad change can spend years of accumulated trust. Turn runbook knowledge into controlled, reusable leverage, so the desk's speed survives technician turnover instead of walking out with it. And adopt AI while protecting the thing that makes the whole model possible: the client's decision to hand over the keys. Notice what is not on the list: adopt AI. Adoption is a means. The moment it becomes the goal, the program starts optimizing for usage instead of for the desk.

The trap

The trap is measuring the means. Assistants deployed, suggestions accepted, scripts generated per technician: activity metrics reward the appearance of change while the operating outcomes sit unmeasured. A desk can hit every adoption target it sets and end the year with the same recurring tickets, the same escalation load, and the same two senior engineers carrying the after-hours phone.

AI adoption is not an outcome. A quieter, better-served client base is.

The checklist

  • Write each of the four outcomes as an operating sentence with a named owner, not a slogan on a slide.
  • Tie every initiative to exactly one outcome. An initiative that maps to none is a hobby.
  • Baseline the desk metrics before the first pilot, or the after will have no before.
  • Retire activity metrics from leadership reporting. Keep them for operations, where they belong.

Where judgment beats the tool

Outcomes conflict at the margin: pushed far enough, capacity presses on quality, and automation presses on control. A dashboard will not arbitrate that tension. Where the firm sets each trade, client by client, is a leadership decision, and it is the one clients experience directly.

Chapter 3

Who has to say yes: the six chairs in the room

The principle

Nothing durable changes in an MSP without the whole operating table, and this program touches every chair at it. The owner asks whether it grows the business without breaking the per-seat model. The service delivery director asks whether the desk gets more reliable or just busier. The vCIOs and account leads ask what happens at the QBR, and whether renewals get safer or riskier. The CTO or automation lead asks whether this becomes one governed platform or thirty ungoverned tools. The security and compliance lead asks what it does to credentials, tenant boundaries, and the trust clients placed in the firm. The CFO asks what it costs, what it returns, and which agreements it rescues. In a smaller MSP, three of those chairs may share one calendar. The questions do not consolidate just because the titles do, and the program has to hold a real answer to all six.

The trap

The trap is the champion-led initiative that answers one chair. An automation push moves fast on borrowed enthusiasm, then dies in a leadership meeting the day the security lead asks the credential question nobody prepared for, and the silence costs the champion a year of credibility. In a business built entirely on being trusted with the keys, the unanswered security question is not a delay. It is a verdict.

The chair you did not brief is the veto you did not see coming.

The checklist

  • Map the six chairs to named people, including the ones who hold the role without the title.
  • Write down each chair's question and the evidence that would satisfy it, before the program is proposed.
  • Brief the security and compliance chair first, not last. In this business that chair holds the sharpest veto.
  • Give the client-facing chairs a voice in what clients will be told, because they are the ones who will be asked.

Where judgment beats the tool

An org chart names the titles. It does not reveal whose no actually ends a program in your firm, or whose quiet yes moves the undecided. Reading the real decision structure of your own shop is judgment, and no tool has ever held it.

Chapter 4 · Flagship workflow

Ticket-to-Knowledge: stop re-solving what the desk already solved

The principle

Ticket-to-Knowledge runs from intake to documented resolution: triage, runbook retrieval, a draft fix grounded in what the firm already knows, escalation that carries its evidence, and a write-back that turns the resolution into firm property. It is the anchor workflow because it is where the firm's own knowledge does the most work: making what the best engineers already know available to everyone else is exactly what a runbook corpus is for. No public study measures this on an MSP desk, so take the supporting pattern for what it is. AI assistance tends to narrow the gap between less and more experienced workers on supported tasks (Noy and Zhang; Dell'Acqua et al.; Brynjolfsson, Li and Raymond; a cross-study inference, not an MSP measurement). That is a tier-one leverage story, and it raises the stakes on supervision rather than lowering them.

The workflow, stage by stage

  • Intake and triage: the ticket arrives enriched with client, environment, and the history of this exact issue, so routing takes seconds instead of archaeology.
  • Runbook retrieval: the relevant runbook and past resolutions surface with provenance, instead of living in whichever senior engineer happens to remember them.
  • Draft resolution: tier one works from the firm's own documented fixes for this client's environment, verifying each step against the live system before acting.
  • Escalation with evidence: what was tried and what the environment showed travels with the ticket, so tier three starts from findings instead of from scratch.
  • Documentation write-back: the resolution becomes a runbook entry as part of closing the ticket, so the knowledge survives the engineer who found it.

The trap

The trap is retrieval over a stale, uncurated pile. Point an assistant at an unmanaged documentation share and it will faithfully suggest last year's procedure against this year's environment. On a writing desk that is a weak draft. On an MSP desk it is a wrong change proposed against a live client system, delivered with confidence to the least experienced person in the room. The corpus is only leverage when someone owns it, reviews it, and retires what is no longer true.

Tribal knowledge is a single point of failure with a notice period.

The checklist

  • Curate the corpus: runbooks with owners and review dates, not the whole documentation share.
  • Wall retrieval by client, so one tenant's environment details never surface in another tenant's ticket.
  • Make the write-back part of closing the ticket, not a virtue reserved for quiet weeks.
  • Keep suggestions as suggestions: the engineer verifies against the live environment, and consequential actions go through the change gate.
  • Track first-touch resolution and escalation rate together, so tier-one leverage is real and not deferred cost.

Where judgment beats the tool

Retrieval can surface the fix that worked last time. It cannot know that this client's environment diverged from the runbook two migrations ago, or that the symptom matches while the cause does not. The engineer's read of a live environment is the difference between a resolution and an incident, and it is not in the corpus.

Open →
Chapter 5 · Flagship workflow

Service-to-QBR Insight: the meeting where contracts renew

The principle

The QBR is the strangest meeting in the MSP model: the one hour in the quarter when the client actually sees the service they have been paying for, and the meeting where the contract quietly renews or quietly begins to die. Most QBR packs are assembled the night before from ticketing exports and memory, which is why so many of them defend the past instead of selling the future. Service-to-QBR Insight rebuilds the pack from live service data: ticket and SLA synthesis, incident narratives with root cause and remediation, per-client trend and risk signals, and roadmap recommendations tied to observed evidence, all drafted for the vCIO's review. The pack stops being a performance and becomes an instrument, and the meeting stops being about last quarter's numbers and starts being about the client's next year, which is the conversation renewals are actually made of.

The workflow, stage by stage

  • Service synthesis: tickets, SLA attainment, and incident history assembled from the service stack, every figure drillable to the underlying records.
  • Incident narrative: what happened, why, and what changed so it does not recur, drafted for review by someone who lived it.
  • Trend and risk signals: recurring tickets, aging assets, and looming exposures surfaced per client, before the client discovers them first.
  • Roadmap drafting: recommendations tied to observed problems and priced from the catalog, so the roadmap reads as evidence rather than as selling.
  • The meeting: the vCIO walks in with instruments and spends the hour on the client's business instead of on defending numbers.

The trap

The trap is synthesis that launders the quarter. A model summarizing service data will happily smooth a rough quarter into confident prose, and the client across the table lived that quarter. They remember the outage, the migration that ran long, the ticket that bounced three times. A polished pack that contradicts the client's lived experience does not reassure them. It teaches them that the reporting cannot be trusted, and the renewal conversation inherits the lesson.

The client remembers the outage. If the QBR does not, the renewal will.

The checklist

  • Make every number in the pack drillable to tickets, alerts, or incidents in one step.
  • Lead with the quarter the client actually had, including what went wrong and what changed because of it.
  • Show recurring-ticket trends honestly, with the root-cause work attached.
  • Tie every roadmap item to observed evidence from that client's own environment.
  • Keep the vCIO's read before the pack ships. Synthesis prepares the story; it does not get to tell it.

Where judgment beats the tool

Synthesis can produce the pack. It cannot read the room, decide which finding leads, judge when to concede a bad quarter plainly, or sense that this renewal needs a conversation before it needs a slide. The relationship is the product, and the vCIO is the one holding it.

Chapter 6 · Flagship workflow

Quote-to-Contract: price the environment, not the hope

The principle

Quote-to-Contract runs from the first look at a new environment to a signed agreement delivery can actually keep: environment assessment, scoping from the service catalog, agreement drafting from approved templates, senior review of scope, price, and exclusions, and a structured handoff into onboarding. Under a per-seat model this workflow carries a weight it carries nowhere else: the price set here is fixed for the life of the agreement, while the cost is discovered afterward, one ticket at a time. A quote built on a day of assessment and a decade of optimism becomes a leak no amount of desk efficiency can plug. AI helps underneath: assessments summarized from discovery data, quotes assembled from a maintained catalog instead of the memory of the last deal, agreements drafted from templates that carry only the service levels delivery has signed off on. The senior review, whether the promise can be kept and what it should cost, stays human, because it was always the point.

The workflow, stage by stage

  • Environment assessment: the new environment documented from discovery data before the price is set, so the quote reflects reality instead of hope.
  • Catalog scoping: services and service levels drawn from a maintained catalog, with exclusions stated instead of discovered.
  • Agreement drafting: first drafts from approved templates, carrying only promises delivery has agreed it can keep.
  • Senior review: an owner or delivery lead reviews scope, price, and exclusions on every agreement, because the fee is fixed and the risk is not.
  • Handoff to onboarding: what was sold arrives as a structured record, so onboarding builds the runbook from the agreement instead of from recollection.

The trap

The trap is quoting what the conversation can win instead of what the desk can keep. When drafting gets cheap, quotes multiply, and the discipline that used to live in scarcity has to live in the review gate instead. An agreement that underprices a noisy environment does not fail at signature. It fails eighteen months later, in effort reports, after-hours resentment, and a renewal where somebody finally does the math out loud.

An underpriced agreement is not a win. It is a leak with a signature.

The checklist

  • Assess before pricing: no quote leaves without an environment reality check behind it.
  • Quote from the catalog, so every promise maps to a service delivery has defined and costed.
  • State exclusions in the agreement, because the ones left unstated get delivered free forever.
  • Hold senior review of scope, price, and exclusions on every agreement, not just the large ones.
  • Hand off every signature as a structured record: what was promised, to whom, at what service level.

Where judgment beats the tool

The catalog can price the services. It cannot price the client: the environment tidied up for the assessment, the culture that files a ticket for everything, the growth that will double the seats by renewal. Reading the risk in a client is the difference between a book of agreements and a book of liabilities, and it is a senior call.

Chapter 7

Onboarding, escalation tiers, and change approval

The principle

Three operating disciplines decide whether growth is safe. Onboarding is where the economics of a client are locked in: the environment documented, credentials captured into controlled storage and nowhere else, runbooks drafted while discovery is fresh, and a client that reaches steady state in days instead of quarters. Escalation tiers are the leverage structure of the desk: tier one resolves what it should, so senior engineers see only what genuinely needs them, and every escalation travels with its history, what was tried and what the environment showed. Change approval is the discipline that makes the other two safe: consequential changes to a client environment pass a gate where an accountable engineer approves, no matter who or what proposed the change. AI assists all three, onboarding documentation drafted from discovery data, escalation summaries assembled from ticket history, change requests written up with their evidence attached, and the approvals stay human, and stay named.

The trap

The trap is onboarding as heroics. The new client is stood up by the best engineer working nights, the documentation is promised for later, and later never comes. Every ticket that client ever files now pays interest on that debt: the desk asks the same engineer, the escalation path routes around the missing runbook, and when that engineer resigns, the client is effectively onboarded again, by someone slower, under pressure, with the client watching.

Every undocumented environment keeps a senior engineer on permanent call.

The checklist

  • Make documentation part of the definition of onboarding done, not a follow-up task.
  • Capture credentials into controlled storage at intake, and nowhere else, ever.
  • Publish the escalation ladder, and require every escalation to carry what was tried and what the environment showed.
  • Route every consequential change through the approval gate, including the ones proposed by automation.
  • Measure onboarding in days from contract to steady state, and read it per client.

Where judgment beats the tool

A checklist can define a consequential change. It cannot weigh one: whether this patch, on this cluster, during this client's quarter-close, is routine or reckless. The change gate works because a human with context stands in it, and context is precisely what does not automate.

Chapter 8

Where the risk lives: credentials, tenants, and client trust

The principle

An MSP is trusted in a way almost no other supplier is: the client handed over the keys. Admin credentials, remote access to every endpoint, standing permission to change things. That trust is the entire franchise, and AI touches it in four places. Credentials: secrets pasted into prompts, tools, and logs that were never approved to hold them. Privileged action: the quality risk is documented rather than hypothetical, since a field experiment with consultants at one elite firm has been reported to show professional performance degrading when AI is relied on for tasks outside its capability frontier (Dell'Acqua et al., 2023; consultants, not engineers), and where that frontier sits shifts with the model and the domain, which is why review stays a standing control rather than a launch-week check. On an MSP desk the outside-frontier failure is not a weak paragraph, it is a plausible-sounding runbook action or security change applied to a live environment without change control. Multi-tenant confidentiality: one client's environment details surfacing in another client's ticket, at machine speed. And disclosure: clients and their auditors do ask how the MSP uses AI inside their environment, sometimes in the security questionnaire, and the firm needs an answer it is proud of before the question arrives (a practice observation; the research library holds no measurement of how common the question is). One rule answers all four, and it holds at every maturity stage: an accountable engineer stays on every consequential action, and AI never holds standing permission to change a client environment.

The trap

The trap is treating security posture as a tooling setting instead of an operating boundary. A tenant toggle does not know which credentials a prompt contains, what this client's compliance overlay prohibits, or that two of your clients are competitors whose environments must never share a context window. General policy plus default settings is how an MSP ends up technically compliant and actually exposed.

Saved hours are invisible. A credential in the wrong prompt is not.

The checklist

  • Vault credentials so they never appear in prompts, outputs, or logs, and audit for the day one does.
  • Wall AI context by client, mirroring the tenant isolation the rest of the stack already enforces.
  • Deny AI standing write access to client environments; every consequential action passes a human gate.
  • Check each client's AI posture against their contract and compliance overlay, because regimes differ.
  • Write the disclosure answer before the questionnaire asks, and rehearse the incident path before the miss.

Where judgment beats the tool

A control framework sets the floor. It cannot weigh whether this client, this environment, this compliance overlay justifies more caution than the standard, or which client deserves a proactive conversation rather than a policy reference. Proportionality is a judgment about relationships and consequences, and it belongs to the people who answer for them.

Chapter 9

The maturity path: baseline, AI-enabled, selectively AI-native

The principle

There are three honest stages, and they belong to workflows, not to firms. Baseline is where most desks live today: runbooks in heads, QBRs by hand, tickets resolved by whoever remembers. It is not failure; it is a leverage ceiling. AI-enabled is the working middle: the same workflow with retrieval, drafting, and synthesis assisting inside governed boundaries, and every gate exactly where it was. Most workflows should live here. Selectively AI-native is the far stage: a workflow redesigned around governed AI because its volume and structure justify it, documentation write-back and QBR data assembly are the natural first candidates, with humans owning judgment and approval by design. One boundary does not move with maturity: a change to a client environment never loses its human gate, at any stage, in any workflow. The operative word is selectively. The path is walked one workflow at a time, on evidence.

The trap

The trap is the platform project that tries to make the whole desk native at once. It burns a year on tooling debates, unsettles the engineers whose buy-in it needs, and usually dies in the gap between a vendor demo and a Tuesday night incident. The opposite trap is subtler: staying enabled forever out of comfort, when the documentation workflow earned its redesign a year ago and the desk is still paying the manual tax.

Workflows earn native status. Firms do not.

The checklist

  • Place each significant workflow on the path separately. The desk does not have a maturity level; its workflows do.
  • Earn native with evidence from enabled: volumes, error rates, and gate outcomes, not vendor roadmaps.
  • Keep the human gate on client-environment changes at every stage. What changes is where drafting and assembly happen, never who approves.
  • Revisit placements quarterly. Workflows earn promotion, and some earn demotion.

Where judgment beats the tool

A maturity model can place a workflow on the path. It cannot decide whether the redesign is worth the disruption this quarter, on this desk, with this bench. Sequencing is strategy, and strategy belongs to the owners.

Open →
Chapter 10

Governance and measurement: rails under the desk

The principle

Governance here is a small set of operating patterns, each of which buys speed as well as safety. Client boundaries decide what AI can reach, mirroring the tenant isolation the stack already enforces. Role-based access mirrors the tier structure. Approved sources decide what suggestions are built from: the curated runbook corpus and the service stack, not the open internet's best guess. Credential vaulting keeps secrets out of prompts. Human change approval stands in front of every consequential action. Traceability rides along, so every AI-assisted resolution shows its sources, which is what makes review fast. Retention rules say what happens to prompts and outputs when a client offboards. Firms that want an external scaffold have two honest options: NIST's AI Risk Management Framework (January 2023) and its Generative AI Profile (July 2024), a voluntary framework for trustworthy AI, and ISO/IEC 42001:2023, the first international AI management system standard, whose certification weight may exceed what a mid-market MSP needs (NIST AI RMF; ISO/IEC 42001). Then measurement, because governance without measurement is theater with paperwork. Six numbers tell you whether the machine is working: seats or endpoints per engineer, first-touch resolution at tier one, mean time to resolution by ticket class, recurring-ticket rate per client, onboarding days from contract to steady state, and effort versus agreement value by client. Baseline them before the first pilot. The adoption surveys are the cautionary tale: 2025 self-reports show regular AI use at 88 percent of organizations while only 39 percent report enterprise-level EBIT impact, and most experiments never fully scale (McKinsey State of AI 2025; Deloitte generative-AI survey 2025). Access is not value capture. Measurement is the difference. And the number to distrust most is the industry benchmark, because for MSP margin, utilization, and seat economics no approved primary source exists in this library. The honest baseline is your own.

The trap

The trap is vendor math wearing a dashboard: tickets deflected that nobody verified, hours saved that nobody redeployed, a benchmark quoted from a webinar with no methods behind it. A leadership team hears one inflated number and discounts the entire program, including the parts that were true. The governance twin of that trap is the binder: a policy nobody opens, an approval queue that takes a week, and a desk that quietly routes around both, leaving the firm with the risk and the bureaucracy at once.

If the governed path is slower than the shadow one, the desk has already chosen.

The checklist

  • Build controls into the path of work as defaults, not into a binder as clauses.
  • Price every control by what it buys: speed gained or risk retired. A control that buys neither is friction wearing a badge.
  • Baseline the six desk numbers before the first pilot, and name one owner per metric with its source of record.
  • Name where freed hours go: more seats per engineer, root-cause closure, a sustainable after-hours rotation. Unassigned capacity evaporates.
  • Check quarterly for the workflows people route around. An evaded control has already failed, whatever the policy says.

Where judgment beats the tool

The measurement can prove the desk got faster. It cannot decide whether the freed capacity becomes growth, margin, or a humane on-call rotation, and it cannot make leadership enforce the choice. Value realization is a leadership act. The tools only make it visible.

Open →Open →
A 90-day way in

You do not transform the desk. You run one honest test. Pick two workflows, one from the service desk and one from the account side, Ticket-to-Knowledge and Service-to-QBR Insight are the natural pair, and spend ninety days proving what governed leverage does to them.

  • Days 1 to 30: baseline the two workflows honestly, set the client boundaries and credential rules, name the accountable owners, and write the quality criteria and change gates where engineers will see them.
  • Days 31 to 60: run the governed pilots with escalation tiers and change approval unchanged, instrument the desk KPIs, and hold a short weekly read of what the numbers and the engineers are saying.
  • Days 61 to 90: read the results against the baseline, codify what worked into runbooks and defaults, retire what did not, and brief leadership on evidence instead of enthusiasm.

The order matters more than the speed. An MSP that baselines, bounds, pilots, and codifies in ninety days knows something true about its own desk, and it has earned the right to pick the next two workflows. That is how selective becomes cumulative.

Sources cited, with their limits. Brynjolfsson, Li & Raymond (NBER w31161, 2023): a deployment of a generative AI assistant in a customer-support setting, single-tenant software support chat with many agents outside the U.S., and not multi-tenant MSP work with SLAs and privileged access; cited in this playbook only as one of the studies behind the experience-gap pattern, and its support-desk productivity findings are not carried across to MSP work here. Dell'Acqua et al., "Navigating the Jagged Technological Frontier" (2023, SSRN): a field experiment with consultants at one elite firm, reporting that quality declined when AI was relied on for tasks outside its capability frontier; consultants, not engineers, and where the frontier sits shifts with the model and the domain. The experience-gap pattern (Noy & Zhang, Science, 2023; Dell'Acqua et al., 2023; Brynjolfsson, Li & Raymond, 2023) is a cross-study inference, not an MSP onboarding measurement. McKinsey State of AI (2025) and Deloitte generative-AI survey (2025): self-reported, large-enterprise-skewed adoption surveys, not MSP financials. NIST AI RMF 1.0 (January 2023) and the Generative AI Profile, NIST.AI.600-1 (July 2024): voluntary frameworks, not mandates and not a liability safe harbor. ISO/IEC 42001:2023: the standard exists and states requirements; certification cost and complexity may exceed a mid-market MSP's capacity. And the gap, stated plainly: no approved primary source in this library covers MSP margin, utilization, or seat economics, which is why this playbook teaches seat economics qualitatively and keeps insisting on your own baseline. What this playbook is not: it does not select tools, it does not predict your results, and it does not replace the accountable engineer, who stays on every consequential action; AI never holds standing permission to change a client environment. The workflow lens is shared with the other industry playbooks in this library; the treatment is specific to managed service providers.
Start with the Workflow Leverage & Trust Assessment →

Part of Governed AI Operations for Professional Services.