What Does McKinsey's 2026 Data Actually Say About AI Agents and Profit?
McKinsey's State of AI 2026 survey, fielded from 4 May to 8 June 2026 across 1,719 respondents, found that 40 percent of organisations above US$1 billion in revenue are now scaling AI agents, up from 27 percent a year earlier. Yet only 37 percent report any EBIT impact from AI, essentially unchanged from 2025.
Two numbers moved in opposite directions. Agent deployment at large enterprises rose by thirteen percentage points in twelve months. The share of organisations able to trace any earnings impact to AI stayed flat.
The share McKinsey classifies as high performers, those capturing meaningful enterprise-wide value, held at roughly 6 percent year over year. Deployment accelerated. Measured financial return did not.
This is the single most important pattern for any executive preparing a 2027 AI budget. It says the constraint is no longer access to capable models. The constraint is everything that sits between a working agent and a line on the profit and loss statement.
Why Does Scaling AI Agents Not Automatically Move EBIT?
An agent creates time savings at the task level. EBIT changes only when saved time is converted into reduced cost or increased revenue. Most organisations complete the first step and never structure the second, so productivity appears in employee experience surveys while the cost base and headcount plan remain untouched.
Consider a claims team of forty people. An agent removes twenty minutes per claim from a process that runs six thousand claims a month. That is roughly 2,000 hours saved a quarter.
Those hours are real. But unless the operating plan reassigns that capacity to absorbed volume growth, deferred hiring, or a service the firm previously outsourced, nothing reaches the income statement. The team simply works with more slack.
The three places value leaks before it reaches EBIT
--- Capacity released but never reallocated, so cost per unit falls on paper while total cost stays flat.
--- Pilots that never leave one department, so the fixed cost of the platform is spread across a fraction of the volume that would justify it.
--- Savings that are real but unattributed, because no baseline was captured before deployment and finance cannot separate AI effects from seasonal or headcount effects.
The third leak is the most common and the most damaging, because it converts a genuine win into an unprovable claim at budget time.
What Is the Four-Layer Diagnostic Framework for Finding the Leak?
The framework tests four layers in sequence: capability, adoption, reallocation, and attribution. Value must survive all four to appear in EBIT. Diagnosing in order matters, because a failure at layer one produces symptoms that look identical to a failure at layer four, and the remedies are completely different.
Layer 1 — Capability. Does the agent complete the task at an accuracy the business will accept without a human re-checking every output? If a reviewer still reads all of it, the agent has added a step rather than removed one. Test this with a sampled accuracy audit against a human baseline, not with a vendor demo.
Layer 2 — Adoption. What proportion of eligible transactions actually route through the agent? A tool used on 15 percent of cases cannot move a cost line no matter how good it is. Adoption is measured in transaction share, not licence count, and licence count is the metric most dashboards report.
Layer 3 — Reallocation. Has the released capacity been assigned a destination in the operating plan? This is a management decision, not a technology one, and it is the layer most programmes never formally reach. Without it, saved hours dissipate.
Layer 4 — Attribution. Can finance isolate the effect? This requires a pre-deployment baseline on the specific metric, a comparison group or comparison period, and an agreed rule for what counts. Attribution designed after launch is nearly always contested.
Run the layers in order. The first one that fails is the one to fix, and fixing a later layer while an earlier one is broken produces no result at all.
How Does This Framework Apply to a Hong Kong Enterprise?
Take a Hong Kong insurance group with 300 staff that deployed an underwriting-support agent nine months ago. Accuracy tested well, licences went to 120 users, and the executive sponsor reported the pilot as a success. Twelve months on, the combined ratio has not moved. Each layer explains a different part of why.
At layer two, transaction share told the real story. Of the cases the agent was designed for, 22 percent went through it. Senior underwriters kept working the way they always had, and the agent absorbed the simpler files that were never the cost problem.
At layer three, the operations director had no mandate to change the staffing plan. Released hours went into longer file reviews, which is a quality gain the firm never asked for and cannot bank.
At layer four, nobody had recorded cycle time or cost per policy in the quarter before launch. When the CFO asked what the agent had returned, the answer was an anecdote.
The Hong Kong dimension sharpens this. Mid-market firms here run lean, so there is rarely a transformation office to own layers three and four. Those layers default to nobody, and a technically successful deployment produces a financially invisible result. The Office of the Privacy Commissioner for Personal Data reported in May 2026 that of 60 organisations it examined, 95 percent were using AI in daily operations and more than half ran three or more AI systems. Usage is widespread. Structured measurement is not.
What Changed in Build Versus Buy This Year?
McKinsey's 2026 survey found that 32 percent of respondents said their organisation decided against purchasing at least one software product or feature because the functionality could be built internally using agentic coding tools. That is a structural shift in how enterprise software budgets get set, and it arrived faster than most procurement policies did.
Around 20 percent of organisations overall are scaling agentic coding tools, rising to 31 percent among large enterprises. The capability gap between big and small organisations is widening in the same direction as the agent gap.
For a department head, the practical consequence is that "we will just build it" is now a credible answer in a vendor negotiation, which changes pricing leverage. It is also a trap, because the build decision moves an ongoing maintenance liability onto a team that was not resourced for it.
The discipline is to apply the same four layers to an internal build. A tool your own engineers ship still has to clear adoption, reallocation, and attribution. Building it does not exempt it from proving a return. If anything, the internal build is harder to measure, because its cost is buried in salaries rather than visible in an invoice.
Related reading: what enterprise AI agent platforms actually cost and how outcome-based AI pricing changes the contract.
How Should You Measure AI Impact So the CFO Can See It?
Pick one metric that already appears in a management report the CFO reads, capture ninety days of baseline before deployment, define in writing what counts as an AI-attributable change, and report against it monthly. A metric invented for the AI programme will be discounted. A metric the CFO already tracks will not.
What a defensible measurement design contains
--- One primary metric drawn from an existing management report, such as cost per transaction, cycle time, or first-contact resolution.
--- A baseline period of at least one quarter recorded before the agent goes live, including the seasonal comparison from the prior year.
--- A control group or control period, so the effect can be separated from other changes happening at the same time.
--- A written attribution rule agreed with finance before launch, naming what the programme will and will not claim.
--- A named owner in the business, not in IT, accountable for the reallocation decision at layer three.
The order matters. Agreeing the attribution rule with finance before launch is what converts a later result from a debate into a number. Teams that skip this step are not being careless. They are usually moving fast under sponsor pressure, and the measurement work feels like a delay. It costs about two weeks and it is the difference between a renewed budget and a cancelled one.
What Goes Wrong When Organisations Skip the Diagnosis?
The common failure is treating flat EBIT as evidence that the technology does not work, then switching vendors. The new platform inherits the same broken layers, produces the same flat result, and the organisation concludes AI is overhyped, having never tested where the value was actually lost.
Four pitfalls worth naming
--- Reporting licence counts as adoption. Seats issued and transactions processed are different numbers, and only one of them relates to cost.
--- Letting the pilot succeed on easy cases. If the agent handles the simplest 20 percent of volume, it cannot touch the workload that drives cost.
--- Leaving reallocation unowned. Technology teams cannot change a staffing plan, and business owners are rarely briefed that they need to.
--- Rebuilding the baseline retrospectively. Numbers reconstructed after the fact are always challenged in a budget review, and they usually lose.
There is also a credibility cost that compounds. A department head who presents an unprovable AI result once will be asked harder questions the next time. The value of getting layer four right is partly financial and partly political.
Conclusion: The Gap Is Managerial, Not Technical
The 2026 evidence points to one conclusion. Model capability is no longer the binding constraint for most enterprises. The binding constraint is the chain that runs from a working agent to a reallocated hour to an attributable number, and that chain is owned by management rather than by engineering.
Diagnose in order. Capability, then adoption, then reallocation, then attribution. Fix the first layer that fails, and do not spend on the next platform until you know which layer broke.
We understand AI. We understand you. With UD by your side, AI never feels cold. Twenty-eight years of working alongside Hong Kong enterprises has taught us that the hardest part of this work is rarely the model. It is the quiet organisational plumbing between a good pilot and a number the board will accept.
Reviewed by the UD enterprise AI team, Hong Kong.
Find Out Which Layer Is Costing You
Now that you have the framework, the next step is establishing where your own organisation sits across the four layers. Start with a structured readiness view, then decide what to measure. We'll walk you through every step, from AI readiness assessment to vendor selection, deployment, and performance tracking, with twenty-eight years of Hong Kong enterprise experience behind it.