Gartner asked 186 IT leaders about Copilot. Two of the bars came out the same length, and that is the main story.
There is a chart in Gartner’s 2026 Microsoft 365 and Copilot survey that ought to be laminated and pinned to a wall or a fridge in the office. It shows two statements and how strongly IT leaders agreed with each.
- Microsoft 365 Copilot has transformed how our employees work. Forty-eight per cent agreed.
- Copilot Chat — the free version — is sufficient for many of our employees’ needs. Forty-seven per cent agreed.
The instinct is to read this as a contradiction, and Gartner’s headline nudges you gently that way. I would argue it is not. It is the same finding stated twice, and the second statement is the invoice for the first.
(These are separate questions on separate bases, so the overlap is invisible. The marginals are enough.)
Where the value is, and where the money is
Content generation such as drafting, editing, summarising, building the slides nobody reads, delivers moderate or high value for 84% of respondents. Knowledge search across the corporate sludge pile: 79%. Top two by a distance.
Copilot Studio, the automation layer where the story stops being about typing faster and starts being about operational change: 53% report no or minimal value.
So the value sits in generation and retrieval, the two capabilities any competent chat interface has offered since 2023, and which the free tier substantially covers. The premium delta lives in Studio, in agents, in deep grounding across tenanted data. That is precisely where the curve falls off a cliff. You cannot charge a premium for a commodity because you bundled it attractively, and fluent text generation stopped being scarce some time ago.
Which reframes the headline barrier. Forty-eight per cent name difficulty proving value or ROI as a top-three obstacle, tying with data exposure for first place. At first glance this seems to be more of a measurement problem: Missing telemetry, missing baselines. In other words, you can buy the maturity framework, run the enablement programme, and the value may be there, but it is not being measured.
However, when half your customers independently volunteer that the free product would do, the difficulty proving in value is not a result poor measurement. It is a measurement returning the correct answer. Thin uplift is hard to evidence because there is not much of it there. Weak business cases fail loudly; marginal ones fail quietly, in the renewal meeting, eighteen months out. Which is only the latest evidence that the binding constraint is almost never the technology.
Tied at the top with ROI: concerns about data exposure and oversharing. Copilot did not create this exposure. The salary spreadsheet in a badly permissioned Teams site, the redundancy notes on a OneDrive with access inherited from a 2019 project, the acquisition folder nobody locked down before the intern rotated off… all of it was already reachable. It was merely unfindable, and organisations quietly reclassified unfindability as security. Copilot’s contribution was to index it and hand every employee a natural-language interface.
Chekhov’s gun, on the wall since about 2014, has now been picked up by something that can read. And you know what happens when you show it in the first act…
We need to consider who owns the remediation: Frame it as a Copilot problem and you buy a Copilot control. Frame it correctly — accumulated access-control debt that AI has made legible — and it becomes an information governance programme with a data owner, a retention policy and an unglamorous multi-year runway. The second framing is right more often than not, and considerably less fun to present to a board.
The tribble problem, and the moat inside it

Eighty per cent say additional governance controls are required before deploying Copilot agents widely. Sixty-eight per cent worry about agent sprawl. Sixty-six per cent say their own people lack the literacy to use agents well.
Agents are tribbles: individually charming, trivially reproducible, and by episode’s end they are in the grain stores and the ventilation ducts. Any platform that makes creation frictionless and leaves lifecycle management as an exercise for the reader will produce sprawl — which is why OpenAI put Codex on your phone and called the leash a feature.
Note what the governance worry actually is. Nobody fears an agent turning malicious. They fear an agent doing exactly what it was told, competently, against a permission model nobody audited — obedience, not rebellion. That 80% is not timidity. It is arithmetic on retrieval plus permissions plus action.
But the revealing number is a comparison. Pre-built agents from third parties are blocked outright by 42% of organisations. Pre-built agents from Microsoft: 20%. Same architecture, same tenant, same permission model, same class of risk. Less than half the blocking rate. There is no security-engineering account of that gap. The difference is contractual: Microsoft’s agents arrive inside an existing agreement, an existing indemnity position, an existing relationship with someone in procurement who answers the phone. Trust tracks the invoice, not the threat model — as it usually does. If you are selling agents into enterprises, that 22-point spread is your competitive problem, and no amount of SOC 2 will close it.
Controlled dependence, at scale
Seventy-two per cent are complementing Microsoft 365 with other tools; 69% of those reach for enterprise AI specifically. They run an average of 2.8 third-party products simply to manage and govern the Microsoft estate. What about the share planning to replace Microsoft 365 functionality? Eleven per cent.
Nobody is leaving. Everybody is hedging at the edges. The scaffolding erected around the platform is now its own budget line, and every plank deepens the dependency it was bought to soften. The moat is not the product; the moat is the exit cost, and the third-party ecosystem is enlarging it while marketing itself as the alternative.
Two footnotes, offered for their comic value. Only 25% agree employees understand and regularly use newer Copilot features — including model choice with Claude; Microsoft has gone to the trouble of shipping a competitor’s frontier model inside its own product and three-quarters of the market has not noticed. And only 20% are likely to buy the new E7 licence against 64% unlikely: when the complaint is that the current tier struggles to prove its worth, the answer arrives as a more expensive tier. You asked for value. Here is more SKU.
What I would actually do
Stop proving ROI in aggregate. Whole-population productivity measurement is a research project, not a business case. Instrument two or three workflows with real cycle-time baselines; let the rest ride on licence economics rather than a benefits claim you cannot defend.
Price the free tier into your renewal model. You almost certainly have a paid cohort whose usage is indistinguishable from the free one. That is a negotiation asset, and it expires the moment your account team finds it first.
Reclassify oversharing as a data programme. Named owner, retention policy, sensitivity labelling, permission remediation with a burn-down chart. Boring, unavoidable, and best done before agents start acting on what they retrieve rather than merely summarising it.
Treat the literacy gap as a leading indicator, not a training need. Enablement decks will not fix it, because the scarce skill is not prompting but verification. Judging whether the output is right is the durable capability, and nobody is building it. Which is the question I keep returning to whenever a tool offers to check its own homework: who checks the chemistry?
A note on the sample
186 IT leaders from Gartner’s own panel and conference outreach, March to May 2026. Self-selected, client-skewed, weighted towards organisations already committed to Microsoft. Gartner says so in their disclaimer and they are right to. Several deployment figures are also calculated among respondents already at least planning to pilot, which removes the abstainers and flatters the maturity picture.
Also: the methodology reports 186 participants, then says 187 completed, then splits the sample 115 plus 71. Which is 186. Somebody’s spreadsheet had a rough afternoon, and I say that with the fondness of a man who has had several.
The signal survives the arithmetic. Half the market thinks the free version would do. That is not a Copilot problem or a measurement problem. It is a pricing problem, and pricing problems resolve themselves at renewal whether or not anyone built the framework first.
Related reading
- Competitive Advantage in the AI Era — why the barrier is rarely the technology
- Houdini in the Sandbox: OpenAI’s Model Goes “Rogue” — obedience, not rebellion
- Claude Science: Who Checks the Chemistry? — verification as the scarce skill