A field playbook · IT consulting

Deliver More Projects with Governed AI

A field playbook for IT consulting and systems-integration firms to raise delivery throughput without diluting engineering quality.

First edition · 2026 · jimiige.com
Jimi Ige · AI Governance, Deployment and Adoption
Why this, and how to read it

Most IT consulting firms grow the same way: win more projects, hire ahead of the backlog, and push utilization until the bench disappears. It works until the scarce thing runs out, and the scarce thing is never developers or laptops. It is the senior engineers and architects who can scope a solution, price the estimate, review the work, and stand behind what ships. Meanwhile the margin was never decided in delivery at all. It was decided at the estimate, weeks before the first commit, and it dies quietly on the bench between projects. AI has already arrived inside this operating model, in generated code and drafted responses and summarized status, whether the firm decided anything about it or not.

This playbook is about operations, not technology. The product is a firm that promises what delivery can keep, builds on what it has already built, and reports on the portfolio in numbers a client can challenge and a CFO can trust. AI is the enabling technology. Governance is the reason the result can be trusted inside a client's environment, and the reason the improvement survives its first fixed-bid overrun. It is written for the person who owns the outcome: the leader who has to raise throughput without diluting the engineering quality the firm's name is built on.

A note on method. Every number in this playbook is cited to a named primary source and carries its own caveat, and where the honest evidence is a gap, the gap is stated instead of filled. Two gaps matter here: no public causal study measures AI's return for mid-market IT consulting firms, and no one has measured what AI assistance does to estimate accuracy, the number this industry lives on. No invented clients, no vendor arithmetic, no borrowed payback periods. The pattern is specific enough to test against your own firm, with a 90-day way to run that test.

Capacity
Deliver more projects without proportional hiring, by moving drafting, retrieval, and assembly onto governed, reusable rails.
Outcome 01
Quality
Raise throughput while architecture and code review gates, and the senior engineers accountable at them, stay exactly where they are.
Outcome 02
Trust
Client IP, credentials, and environments protected by design, with provenance the firm can show on request.
Outcome 03

Three commitments, no hockey sticks. Each chapter ends with where judgment beats the tool, because in this business the estimate and the architecture call are the product.

In-frontier
on tasks inside AI's capability frontier, consultants in a field experiment at one elite global firm have been reported completing more work, faster, at higher rated quality than a control group
Dell'Acqua et al., 2023 · one elite firm, experimental consulting tasks, not engineering delivery and not mid-market proof
~40%
less time on occupation-specific professional writing tasks, with quality up about 18 percent, in a preregistered experiment (N≈453)
Noy & Zhang, Science, 2023
Utilization
surveyed billable utilization across professional-services firms has been reported at its lowest level in that benchmark's own history, and below the benchmark's own target
SPI Research, summarized by Deltek · a vendor-sponsored voluntary survey, multi-vertical with no public IT-consulting split, and not an AI effect

Frame these precisely. The strongest causal evidence comes from consulting-style knowledge tasks at one elite firm and from a controlled writing experiment, not from engineering delivery; the utilization benchmark is a vendor-sponsored, multi-vertical survey whose IT-consulting split is not public. None of them proves firm-level margin for a mid-market IT consulting firm, no public study does, and the number this industry most wants, what AI does to estimate accuracy, has not been measured anywhere. The same field experiment also reported quality falling when AI was used on tasks outside its frontier, and where that frontier sits shifts with the model, the stack, and the task, which is why the review-gate chapters exist. Full source notes close the playbook.

Chapter 1

The estimate and the bench: how an IT consulting firm actually earns

The principle

Strip the mystique and an IT consulting firm is a machine for converting engineering hours into promised outcomes. It sells two ways, and the difference runs the firm. On time and materials, the client carries the estimation risk and presses the rate. On a fixed bid, the firm carries it, and the project's margin is decided before the first commit, in the estimate: scope read right or wrong, assumptions priced or missed, exclusions written or forgotten. After the signature, only three things move the number. Reuse, because a solution built on proven components costs less than one built from scratch. The review gates, because a defect caught at design review costs a conversation and the same defect caught in production costs the margin. And the bench, because engineers between projects are pure cost, and every week of bench is a week some project's profit quietly pays for. The sector backdrop is not generous: surveyed billable utilization across professional-services firms has been reported at its lowest level in that benchmark's history and below its own target (SPI / Deltek, sponsored benchmark; a multi-vertical voluntary survey whose IT-consulting split is not public, and not an AI effect). Above all of it sits the binding constraint: the senior engineers and architects who can scope, estimate, review, and stand behind what ships do not scale by hiring, because the market takes years to certify one.

The trap

The trap is buying coding assistants while leaving the operating model alone. Code arrives faster, and nothing else changes: estimates are still guesses dressed as line items, SOWs still promise scope nobody reviewed, the same integration gets rebuilt because nobody can find the last one, and the senior engineers are still the bottleneck at every gate. The licenses were real. The margin never moved, because the margin was never in the typing.

Margin is decided at the estimate, defended at the review gate, and lost on the bench.

The checklist

  • Name the constraint in each practice: winning work, scoping it, building it, or reviewing it. They are different problems with different fixes.
  • Follow one week of senior-engineer hours and mark which of them only that engineer could have spent.
  • Pull the estimate-to-actual variance on the last eight fixed bids, and read where the misses actually came from.
  • Count how often the firm rebuilds something it has already shipped: integrations, modules, runbooks, solution outlines.
  • Choose workflows to improve, not tools to buy. A tool dropped on an unchanged workflow returns almost nothing.

Where judgment beats the tool

A utilization report shows where the hours went. It cannot say which slow hours were the product. The architect who spends a day interrogating a requirement before pricing it may be the most profitable person in the firm that week, and hurrying that day would cost more than it saves. Deciding which senior hours are the value and which are habit is a call only the leadership can make.

Chapter 2

Four outcomes that matter, and one that does not

The principle

Four outcomes justify this whole program. Deliver more projects without proportional hiring, so revenue can grow faster than headcount. Raise throughput without diluting engineering quality, so the growth does not quietly spend the firm's reputation for work that ships and stays shipped. Turn what the firm has already engineered into controlled, reusable leverage, so its best solutions compound instead of retiring with their authors or expiring with a contract. And adopt AI while protecting client IP, environments, and trust, because a firm that is handed the keys to other people's systems is selling trust before it sells anything else. Notice what is not on the list: adopt AI. Adoption is a means. The moment it becomes the goal, the program starts optimizing for usage instead of for the firm.

The trap

The trap is measuring the means. Licenses issued, seats active, the share of code that was AI-generated: activity metrics reward the appearance of change while the operating outcomes sit unmeasured. A firm can hit every adoption target it sets and end the year with the same estimate variance, the same rebuilt integrations, and the same senior engineers underwater at the same gates.

AI adoption is not an outcome. Projects delivered at their estimate are.

The checklist

  • Write each of the four outcomes as an operating sentence with a named owner, not a slogan on a slide.
  • Tie every initiative to exactly one outcome. An initiative that maps to none is a hobby.
  • Baseline the outcome metrics before the first pilot, or the after will have no before.
  • Retire activity metrics from executive reporting. Keep them for operations, where they belong.

Where judgment beats the tool

Outcomes conflict at the margin: pushed far enough, throughput presses on quality, and reuse presses on fit. A dashboard will not arbitrate that tension. Where the firm sets each trade, bid by bid and client by client, is a leadership decision, and it is the one clients experience as the difference between a partner and a body shop.

Chapter 3

Who has to say yes: the six chairs in the room

The principle

Nothing durable ships in a consultancy without the room agreeing, and this program touches every chair in it. The CEO or managing director asks whether it grows the firm without diluting what the name stands for. The COO or delivery director asks whether delivery gets more predictable or just faster on paper. Practice leads and solution directors ask what happens to the quality of the solutions that carry their names. The CTO asks whether this becomes one governed platform or a different assistant in every project, and carries a question the other chairs do not: a firm that implements AI for clients cannot be seen running it ungoverned at home. The security and risk lead asks what it does to client IP, credentials, and environment access. The CFO asks what it costs, what it returns, and whether the number would survive the firm's own estimation review. Six different questions, and the program has to hold a real answer to all of them.

The trap

The trap is the champion-led initiative that answers one chair. It moves fast on borrowed enthusiasm, then dies in a leadership meeting the day the security lead asks which client environments the pilot's tools could reach, and the silence costs the champion a year of credibility. In a firm that sells engineering discipline, an unanswered governance question is not a detail. It is a work sample.

Six chairs, six vetoes. The program that survives answered every one before the meeting.

The checklist

  • Map the six chairs to named people, including the ones who hold the role without the title.
  • Write down each chair's question and the evidence that would satisfy it, before the program is proposed.
  • Brief the skeptics privately before the leadership meeting, not at it.
  • Give the security and risk lead a genuine design seat. Controls added at the end read as concessions; controls designed in read as competence.

Where judgment beats the tool

An org chart names the titles. It does not reveal whose no actually ends a program in your firm, or which delivery leader's quiet endorsement moves the rest. Reading the real decision structure of your own firm is judgment, and no tool has ever held it.

Chapter 4 · Flagship workflow

Pursuit-to-SOW: promise what delivery can keep

The principle

Pursuit-to-SOW runs from a qualified opportunity to a signed statement of work: the RFP response, the solution outline, the estimate, the SOW draft, and the technical review that stands between the firm and its promises. It is the workflow where margin is decided, because everything delivery will later fight for, scope, assumptions, exclusions, rates, acceptance criteria, is written here, usually under deadline, usually by the same senior people the delivery portfolio is already consuming. The drafting layer is where the causal evidence is strongest: in a preregistered experiment, access to a general AI assistant cut time on professional writing tasks by about 40 percent while raising judged quality about 18 percent (Noy and Zhang, Science, 2023; N≈453; task-level results on writing tasks, not engagements and not engineering estimates). The redesign moves assembly onto governed rails, response structure, solution narratives, comparable past work, so senior attention lands on the three calls that decide the project's economics: what to promise, what to assume, and what to charge.

The workflow, stage by stage

  • RFP response and qualification: what the firm knows about the client, the domain, and its own relevant delivery history arrives assembled with sources, so the bid or no-bid call takes senior minutes instead of senior evenings.
  • Solution outline: architectures start from the firm's reference patterns with provenance, not from a blank whiteboard under deadline.
  • Estimation support: comparable past projects and their actuals surface next to the new scope, so the estimate is argued against evidence. The judgment stays human, and no study has measured what AI does to estimate accuracy, so treat assistance here as assembly, never as an answer.
  • SOW drafting: scope, assumptions, exclusions, and acceptance criteria draft from approved language, so the contract says what the estimate priced.
  • Technical review: a senior engineer reviews what is being promised before it is signed, with the estimate, the assumptions, and the delivery plan in one place.

The trap

The trap is a proposal machine that outruns estimation discipline. When responses get cheap, the temptation is to answer every RFP, and pursuit discipline quietly dies. Worse, fluent drafting makes an unpriced promise read exactly like a priced one: the SOW is beautifully written, and the scope it promises was never estimated by anyone. Volume goes up, win rate goes down, and delivery inherits fixed bids the firm should never have signed.

The most expensive sentence in this business is a fluent promise the estimate never priced.

The checklist

  • Keep the bid decision ahead of the drafting engine. No response starts before the pursuit call is made.
  • Require an estimate behind every scope sentence in the SOW. If nobody priced it, nobody promises it.
  • Build the estimation base from your own actuals, with provenance, and argue every new estimate against it.
  • Hold senior technical review on every SOW before signature. Speed is not a reason to skip the one review that decides margin.
  • Track estimate-to-actual variance by project type, and read it quarterly. It is the only measure of whether this workflow is actually improving.

Where judgment beats the tool

Assembly can put the comparables on the table. It cannot decide whether this client's environment will fight the architecture, whether the unknowns justify a contingency the price can carry, or whether the deal is worth the bench it will consume. The estimate is not a document. It is the firm's judgment about the future, priced, and it does not delegate.

Open →
Chapter 5 · Flagship workflow

Build on What You Built: stop rebuilding what the firm already shipped

The principle

An IT consulting firm's edge is what it has already engineered, and most firms cannot retrieve it. The integration pattern lives in a repo nobody remembers, the reference architecture in a departed architect's slides, the hard-won runbook in a project channel that got archived. Build on What You Built turns that delivery history into working leverage: prior-solution retrieval, reference architectures, component and pattern reuse, and documentation drafting grounded in what the firm has already shipped. Done well, every project starts from the firm's best prior engineering instead of from zero. The strongest causal evidence for the underlying pattern comes from consulting-style knowledge work: in a field experiment at one elite global firm, AI assistance on tasks inside its capability frontier has been reported to increase the volume and speed of completed work and to raise its rated quality (Dell'Acqua et al., 2023; one elite firm and experimental tasks, not engineering delivery and not mid-market proof). The reuse thesis in this chapter does not rest on that experiment. It rests on a plainer claim a firm can test on itself: teams that can retrieve the firm's best prior engineering start from it, and teams that cannot quietly rebuild it.

The workflow, stage by stage

  • Prior-solution retrieval: the relevant repos, designs, and deliverables surface with provenance and contract context, instead of living in whoever happens to remember them.
  • Reference architectures: proven patterns arrive versioned and owned, so solutioning starts from the firm's best, not from memory under deadline.
  • Component and pattern reuse: reusable components carry an owner, a review date, and an IP status, so reuse never means shipping one client's property to another.
  • Documentation drafting: designs, runbooks, and handover documents draft from the delivered system's own artifacts, so documentation stops being the thing the budget runs out before.
  • Write-back at close: what the project taught the firm, components, patterns, estimates against actuals, is filed as part of shipping, so the corpus compounds instead of decaying.

The trap

The trap is reuse without provenance or boundaries. Point an assistant at the firm's whole delivery history and it will faithfully surface code the firm does not own, patterns a contract left behind with a client, and components nobody has reviewed since the engineer who built them left. Generated code sharpens the same problem: a snippet with no provenance record, accepted at speed, is a license and quality question the firm answers later, in someone else's production environment.

A component you cannot trace is a liability you have already shipped.

The checklist

  • Curate the reuse library: reviewed components and reference architectures with owners, review dates, and IP status, not the whole repo history.
  • Wall retrieval by client and contract, so what one engagement built never surfaces in another uninvited.
  • Attach provenance to every reused and generated artifact, so a reviewer can see where each piece came from before it ships.
  • Keep the review gates ahead of every reuse. A proven component in a new context is a new engineering decision.
  • Make the close-out write-back a deliverable with an owner, not an intention.

Where judgment beats the tool

Retrieval can find the architecture from the project that resembles this one. It cannot tell you which parts worked because of that client's data volumes, that team's skills, that year's platform quirks. Knowing what generalizes and what was circumstance is what makes a senior engineer senior, and it is the judgment the reuse library serves rather than replaces.

Chapter 6 · Flagship workflow

Project-to-Portfolio Signal: reporting the firm can run on

The principle

Leadership runs the portfolio on synthesized truth: status, burn, velocity, risk, margin. In most firms that synthesis is assembled by hand on Friday from project-manager memory and optimism, and on fixed bids the distance between reported and real is measured in margin. Project-to-Portfolio Signal rebuilds the picture from live delivery artifacts instead: status synthesized from repos, boards, and timesheets; burn and velocity on one spine across projects; margin by project while the project can still be corrected; risk escalated to a named owner with a date; and client-ready reporting the engagement lead reviews before it ships. The pack stops being a performance and becomes an instrument.

The workflow, stage by stage

  • Status synthesis: the weekly picture assembles from the artifacts delivery already produces, so reporting stops taxing the people doing the work.
  • Burn and velocity: spend against estimate and progress against plan sit on one spine, comparable across projects and readable in one pass.
  • Margin by project: fixed-bid burn reads against the estimate weekly, so an overrun is a correction in week three instead of a write-off in week thirteen.
  • Risk escalation: risks route to a named owner with a date, never to a distribution list.
  • Client-ready reporting: what the client sees is drawn from the same spine the firm runs on, reviewed by the engagement lead before it leaves.

The trap

The trap is synthesis that launders slippage. A model summarizing status reports will smooth a troubled project into confident prose, and a fixed bid can burn quietly under green reporting for a month, which on a fixed bid is the month the margin leaves. The failure is not the summary. It is a reporting chain where no number can be challenged, because no number can be traced.

By the time a fixed bid admits it is red, the margin has already left.

The checklist

  • Give every metric one source of record, and synthesize from sources, not from prior summaries.
  • Make every portfolio number drillable to the artifact behind it, in one step.
  • Read fixed-bid burn against estimate weekly, with the variance visible to someone who can act on it.
  • Escalate risk to a named owner with a date, never to a distribution list.
  • Keep an engagement-lead read on every client-facing report. Synthesis prepares the story; it does not get to tell it.

Where judgment beats the tool

Synthesis can surface the anomaly. It cannot decide which slipping project is worth a senior architect's week, which client call gets made before the number is certain, or when a recovery plan should become a leadership change. Escalation is a judgment about consequences, and it belongs to people with their names on the account.

Chapter 7

Review gates and environment onboarding: where quality is decided

The principle

Two disciplines keep an IT consulting firm's quality bar where clients believe it is, and AI raises the stakes on both. The first is the review gates: design review before the architecture is committed, code review before anything merges, and an accountable senior engineer at each, answering for what ships whoever or whatever produced it. AI moves more work toward those gates faster, and the work arrives fluent, which is precisely the problem. The field evidence has been reported to show quality declining when professionals rely on AI for tasks outside its capability frontier (Dell'Acqua et al., 2023; one elite firm, experimental tasks), and where that frontier sits shifts with the model, the stack, and the task, which is why the gates are standing controls rather than launch-phase ones. The second discipline is environment onboarding: the access, credentials, and security-posture documentation that begin every project. The firm works inside its clients' systems. That standing fact decides what AI may touch: which environments assisted tools can see, which client rules govern them, and where credentials live, which is never in a prompt.

The trap

The trap is review that decays into skimming as volume rises. A gate that approves fluent work unread is not a gate, it is a formality with a senior engineer's name on it, and throughput that outruns review capacity is not capacity. It is exposure, accumulating interest until acceptance testing or production finds it.

The client never asks what generated the code. They ask who reviewed it.

The checklist

  • Name the accountable senior engineer at every design and code gate, before the work arrives, not after.
  • Budget review capacity like the scarce resource it is. When throughput rises, review time is the first thing to protect, not the first to trim.
  • Mark AI-assisted work as such at the gate, with its provenance, so reviewers know what they are reading and review it accordingly.
  • Provision environment access with project staffing and revoke it the same way, on the firm's side and the client's.
  • Write each client's rules for AI in their environment into onboarding before the first assisted commit, and keep credentials out of prompts absolutely.

Where judgment beats the tool

A gate can force the review to happen. It cannot make the architecture call: whether this design survives this client's scale, whether the elegant solution is the maintainable one, whether the generated approach fits the constraint the requirement never wrote down. Review is not proofreading. It is engineering, and it is what the client is paying the senior rate for.

Chapter 8

Where the risk lives: client IP, credentials, and the fixed-bid exposure

The principle

An IT consulting firm holds a strange double trust: it works inside its clients' systems, and its clients' competitors may be its clients too. AI touches that trust in four places. Client IP: what the firm built on one client's budget, what it may carry forward, and what a contract left behind; retrieval and reuse make that boundary easy to cross at machine speed. Credentials and environment access: the firm operates in customer environments under least privilege, and an assisted tool with unclear reach turns one firm's convenience into another organization's incident. Provenance of generated code: what was generated, from what, under which license posture, and who reviewed it; a question clients, auditors, and acquirers do ask, and the honest answer cannot be reconstructed later (a practice observation; the research library holds no measurement of how often it is asked). And fixed-bid quality exposure: on a fixed bid the firm owns the rework, so a defect that clears the gates and fails at acceptance is not a quality statistic. It is margin, paid back with interest. None of these risks argues against the program. All of them argue for running it inside boundaries.

The trap

The trap is treating these as contract language instead of operating controls. The MSA says the client owns the work product; nothing in the toolchain knows that. The security addendum limits environment access; the assistant's context never read it. A boundary that lives in a document and not in the systems is a boundary the firm is trusting everyone to remember, at speed, under deadline, forever.

Inside a client's environment, the firm's AI posture is part of the client's attack surface.

The checklist

  • Wall reuse and retrieval by client and contract, and record the IP status of every reusable component where the reuse happens.
  • Scope assisted tools to the environments the project actually needs, provisioned and revoked with staffing, and keep credentials out of prompts and context absolutely.
  • Keep a provenance record for generated code as it is written: what was generated, reviewed by whom, merged when. It cannot be assembled retroactively.
  • Review each project's AI posture against that client's contract and security addendum, because they differ, and some now speak to AI directly.
  • Stand up an incident path that assumes a miss will eventually happen, and rehearses who tells the client, how fast, and with what in hand.

Where judgment beats the tool

A control framework sets the floor. It cannot weigh whether this client, this environment, or this bid justifies more caution than the standard, and it cannot hold the conversation with the client whose security team wants to walk through the firm's AI use themselves. Proportion and candor are relationship judgments, and they belong to whoever owns the account.

Chapter 9

The maturity path: baseline, AI-enabled, selectively AI-native

The principle

There are three honest stages, and they belong to workflows, not to firms. Baseline is where most delivery work lives today: patterns in senior heads, estimates from experience and hope, documentation whenever the budget allows, status by asking. It is not failure; it is a leverage ceiling. AI-enabled is the working middle: the same workflow with drafting, retrieval, and synthesis assisting inside governed boundaries, and the review gates exactly where they were. Most workflows should live here. Selectively AI-native is the far stage: a workflow redesigned around governed AI because its volume and structure justify it, and documentation drafting, status synthesis, and response assembly are the usual earners, with humans owning judgment and approval by design. Estimation is the deliberate exception: it can be enabled with comparables and assembly, but the judgment at its center is the firm's risk position, and redesigning it around a model would move the firm's economics onto ground nobody has measured. The operative word is selectively. The path is walked one workflow at a time, on evidence.

The trap

The trap is the enterprise transformation that tries to make the whole firm native at once, burns two quarters on platform debates, and dies in the next pipeline dip. The opposite trap is subtler: staying enabled forever out of comfort, while documentation and status synthesis have long since earned the redesign and quietly keep consuming engineer-hours the bids never priced.

Maturity belongs to workflows, not to firms, and the estimate walks the path last.

The checklist

  • Place each significant workflow on the path separately. The firm does not have a maturity level; its workflows do.
  • Earn native with evidence from enabled: volumes, defect rates, and review outcomes, not enthusiasm.
  • Keep judgment and approval human at every stage, including native. What moves is assembly. What never moves is accountability.
  • Hold estimation at enabled until your own estimate-to-actual data argues otherwise. No public study will make this call for you.
  • Revisit placements quarterly. Workflows earn promotion, and some earn demotion.

Where judgment beats the tool

A maturity model can place a workflow on the path. It cannot decide whether the redesign is worth the disruption this year, with this bench, this backlog, and this client mix. Sequencing is strategy, and strategy is what the leadership table is for.

Open →
Chapter 10

Governance that speeds delivery up, and measurement that keeps it honest

The principle

Governance here is a small set of operating patterns, each of which buys speed as well as safety. Client and project boundaries decide what AI can reach. Access mirrors project staffing, provisioned and revoked with it. Approved tools and approved sources decide what work may be built from. Provenance rides along on everything generated or reused, which is what makes the gates fast. Review gates stand before anything merges or ships, with quality criteria written where reviewers work. Named approval points decide who says ship. Retention rules say what happens to prompts, outputs, and context when a project closes, and reuse rules say what the firm may carry forward and what stays behind with the client. For the scaffold, NIST's AI Risk Management Framework (January 2023) and its Generative AI Profile (July 2024) offer a voluntary, vendor-neutral structure to hang these controls on (NIST AI RMF 1.0; NIST.AI.600-1; voluntary, not a mandate and not a liability safe harbor, adapted to the firm's own client obligations). Then the second half, because governance without measurement is hope. Baseline the six numbers this program answers to: estimate-to-actual variance by project type, proposal and SOW cycle time, the share of delivered work built on reviewed reusable components, defect and rework rates caught at gates versus in production, bench time between assignments, and margin by project during delivery. Count program cost once, expect value after an adoption lag, and present ranges the leadership can believe.

The trap

The trap comes as a matched pair. Governance as committee: a policy nobody opens, an approval queue that takes a week, and engineers who route around both, leaving the firm the risk and the bureaucracy at once. And measurement as marketing: developer-productivity theater, hours saved that nobody redeployed, benefits counted in three business cases at once. The adoption surveys are the cautionary tale: enterprise self-reports in 2025 put regular AI use in at least one function at 88 percent of organizations while only 39 percent report enterprise-level earnings impact, and most experiments do not fully scale (McKinsey State of AI 2025; Deloitte generative-AI survey 2025; self-reported, large-enterprise-skewed consultancy surveys, not IT-consulting figures). Access is not value capture. And say the gap plainly: no public causal study measures AI's return for a mid-market IT consulting firm, which is why your own baseline is not optional. It is the evidence.

Hours saved is a vendor's number. Estimate-to-actual variance is yours.

The checklist

  • Build controls into the path of work as defaults, not into a document as clauses.
  • Price every control by what it buys: speed gained or risk retired. A control that buys neither is friction wearing a badge.
  • Baseline every KPI before the pilot that is supposed to move it, with one owner and one source of record per number.
  • Count shared program cost once, model the adoption lag explicitly, and present ranges with the assumptions attached.
  • Name where freed hours land: more projects, a shorter bench, sharper estimates, senior time back at the gates. Unassigned capacity evaporates.

Where judgment beats the tool

Patterns can be adopted; proportion cannot. How much boundary a two-week assessment needs against a two-year platform build, when a control has become theater, and whether freed capacity becomes growth, margin, or a bench that finally gets to breathe: those are calls about consequences, and they belong to the people who answer for them.

Open →Open →
A 90-day way in

You do not transform the firm. You run one honest test. Pick two workflows, one from the pursuit side and one from delivery, and spend ninety days proving what governed leverage does to them.

  • Days 1 to 30: baseline the two workflows honestly, estimate variance and cycle times included, set the client and project boundaries and approved tools, name the accountable engineers at every gate, and write the quality criteria where reviewers will see them.
  • Days 31 to 60: run the governed pilots with the review gates unchanged, instrument the KPIs, and hold a short weekly read of what the numbers and the engineers are saying.
  • Days 61 to 90: read the results against the baseline, codify what worked into defaults and reuse rules, retire what did not, and brief the leadership on evidence instead of enthusiasm.

The order matters more than the speed. A firm that baselines, bounds, pilots, and codifies in ninety days knows something true about itself, and it has earned the right to pick the next two workflows. That is how selective becomes cumulative.

Sources cited, with their limits. Noy & Zhang, Science (2023): a preregistered experiment on occupation-specific writing tasks, N≈453; task-level and evaluator-scored, relevant to the drafting stages of pursuit, not to engineering delivery or estimates. Dell'Acqua et al., "Navigating the Jagged Technological Frontier" (2023, SSRN): a field experiment with consultants at one elite global firm, reporting gains on tasks inside AI's capability frontier and quality damage on tasks outside it; consulting-style knowledge tasks, the closest causal analog for this industry's delivery work, and still not mid-market IT-consulting proof. SPI Research / Deltek professional-services benchmark (2025 performance): a vendor-sponsored voluntary survey across professional-services verticals; the IT-consulting split is not public, and the utilization figure is not caused by AI. McKinsey State of AI (2025) and Deloitte generative-AI survey (2025): self-reported, large-enterprise-skewed adoption surveys. NIST AI RMF 1.0 (January 2023) and Generative AI Profile NIST.AI.600-1 (July 2024): voluntary frameworks, not mandates and not a liability safe harbor. Two gaps are stated in this playbook rather than filled: no public causal ROI evidence exists for mid-market IT consulting firms, and the effect of AI assistance on estimate accuracy is unmeasured anywhere public, which is why this playbook keeps insisting on your own baseline. What this playbook is not: it does not select tools, it does not predict your results, and it does not replace the engineering judgment your clients are paying for. The workflow lens is shared with the other industry playbooks in this library; the treatment is specific to IT consulting.
Start with the Workflow Leverage & Trust Assessment →

Part of Governed AI Operations for Professional Services.