AI performance measurement

Why can’t your AI pilot see the value your people are finding?

Fifty-six percent of CEOs report neither higher revenues nor lower costs from AI over the past year. In the same period, Gallup found that among employees applying AI across seven or more parts of their job, 90 percent say it has improved their productivity. Both findings are credible. The distance between them is a measurement problem wearing a technology costume.

Key takeaways

AI value is discovered by individuals inside the specific texture of their own work, but organizations can only see that value where measurement infrastructure happened to exist beforehand.

  • The pilot is the wrong instrument. Asking people to nominate use cases before they have used a tool produces a forecast, and forecasts are what most AI programs are measuring against.
  • Engagement, not access, separates the workers who gain from those who do not. Two employees with identical tools and identical months of exposure can end up in entirely different places.
  • Training improves prompting and can degrade judgment. In one randomized study, the trained group was both the fastest and the least accurate on work that sat just past the model’s competence.

In 1998, BusinessWeek’s Andy Reinhardt asked Steve Jobs whether Apple had done consumer research on the iMac while developing it. The answer is usually quoted in fragments, and the fragments have done some damage. In full:

No. We have a lot of customers, and we have a lot of research into our installed base. We also watch industry trends pretty carefully. But in the end, for something this complicated, it’s really hard to design products by focus groups. A lot of times, people don’t know what they want until you show it to them.

Read whole, this is not the anti-research manifesto it is usually presented as. Jobs is describing research Apple did do, and then drawing a line: for something sufficiently new, asking people to describe the value of a thing they have not used produces a bad answer, and it is not their fault.

That line is worth recovering, because the method Jobs was warning about is almost exactly the design of the enterprise AI pilot. Assemble a group who have barely touched the tool. Ask them to nominate use cases. Build a business case from those projections. Measure against it. Report the shortfall to a steering committee.

We think this is why so many organizations are simultaneously confident that AI is working and unable to demonstrate it.

The focus group problem

PwC’s 29th Global CEO Survey, covering 4,454 chief executives across 95 countries, found that 56 percent reported neither higher revenues nor lower costs attributable to AI over the preceding twelve months. That figure is widely misquoted as CEOs reporting no benefit at all. It says something more specific and more interesting: asked to attribute a financial movement to one technology, most could not.

Consider what that question demands. It asks the person furthest from the keyboard to isolate the contribution of a single input to twelve months of company-wide financial performance, on a three-point scale where a half-percent saving and a forty-percent saving occupy the same box. A firm with materially faster cycle times and flat revenue lands in the 56 percent. So does the 22 percent who reported that AI increased their costs, a figure that rarely survives into secondary coverage.

The survey is not badly built. It is the same instrument most boards are currently using, and it is structurally incapable of detecting value that is real, distributed across many people, and not yet aggregated into a line item.

What discovery actually looks like

The clearest evidence on how AI value actually arrives comes from a study of 5,172 customer support agents at a Fortune 500 software company, published in the Quarterly Journal of Economics. Resolutions per hour rose 15 percent on average and 30 percent among the least skilled and least experienced agents.

The more instructive finding came from the system’s failures. The vendor’s AI went down periodically, cutting recommendations to some agents and not others, sometimes for minutes and sometimes for hours. Researchers treated these outages as a natural experiment. When an outage struck one month after rollout, affected agents performed at roughly their pre-AI baseline. When one struck after three months, they remained faster than baseline without any AI assistance at all.

Something had transferred into the worker. But not into every worker. Agents who had followed the AI’s suggestions early retained the knowledge gained during outages. Agents who routinely disregarded them showed no improvement, in the authors’ words, “even after prolonged AI access.”

That distinction deserves emphasis, because it is the one most enterprise reporting cannot capture. Access was identical. Tenure was identical. Engagement was not, and engagement is what determined the outcome.

A second study points the same direction from a very different population. Boston Consulting Group ran two randomized experiments across 758 consultants, in work published by Harvard Business School. Afterward, researchers read 244 complete chat logs and found they had to invent vocabulary for the working styles that had emerged. Some consultants delegated whole subtasks to the model. Others interleaved at the sentence level. Nobody had been taught either approach. They appeared anyway, alongside an unassigned repertoire of tactics: assigning the model a persona, pushing back on its output, teaching it by example, requiring it to explain its own reasoning.

Neither study describes a workforce waiting to be handed a use case. Both describe people finding one, unevenly, at different speeds, through repeated contact.

The legibility gap

Here is the pattern we think connects these findings, and the reason the boardroom view and the desk-level view keep failing to reconcile.

Value produced through discovery becomes visible to an organization only where measurement infrastructure already existed for other reasons. Call this the legibility gap: the distance between the value an organization’s people are generating and the portion of it the organization is equipped to see.

Note why the support agent study could detect anything at all. Contact centers have run on average handling time and resolutions per hour for decades. That instrumentation predates the AI by a generation, was built for entirely unrelated purposes, and happened to be pointed at precisely the behavior that changed. The variation between agents who engaged and agents who did not landed on a dashboard someone was already reading every morning.

Now place the same tool into legal review, account management, curriculum design, or grant writing. The same variation almost certainly exists. Almost none of it is counted by the hour. The value is not absent; it is illegible, much as a firm’s institutional knowledge is entirely real and entirely missing from its balance sheet.

The legibility gap explains the apparent contradiction in the PwC data without requiring anyone to be wrong. CEOs are reporting accurately on what their instruments show them. Their instruments were built to detect capital expenditure and headcount, not diffuse improvements distributed across several thousand individual workflows.

Where training backfires

The intuitive remedy is training, and the evidence here should give buyers pause. In the BCG study, one task was deliberately constructed to sit just past the model’s competence. Financial data pointed toward one conclusion; interviews with company insiders pointed toward another, and reaching the right answer required weighing the second against the first. Consultants working without AI answered correctly 84.5 percent of the time. Consultants using AI answered correctly less often, by roughly 19 percentage points. The group that had received prompt engineering training performed worst of all, at 60 percent, and completed the task fastest.

The training worked as designed. It improved prompting, and it built confidence in the output. Confidence is precisely the wrong asset to carry into a question the model cannot answer. The study’s authors raise a similar interpretation, suggesting training may have increased participants’ willingness to defer to the tool.

The practical posture is trust but verify. A capable model behaves like a strong new hire who is occasionally, fluently wrong: fast, articulate, and entirely missing the tribal knowledge that would tell them the spreadsheet is not the whole story. In the BCG case, that is not a metaphor. It is a literal description of the task.

Three groups, three blind spots

The legibility gap looks different depending on where a reader sits, and we think it is worth being specific about who is missing what.

  • Executives are reading a forecast and calling it a result. Pilot ROI is a projection made by people who had not yet found the value, measured against an instrument built for capital allocation. The number is not wrong so much as it is answering a different question.
  • Managers control the variable nobody is measuring. Whether a team member accumulates enough attempts to discover anything depends heavily on whether experimentation is treated as work or as time not spent on deliverables. Few managers recognize this as a decision they are making.
  • Individual contributors are waiting for an assignment that is not coming. The organization can supply tools, permission and hours. It cannot know which parts of a specific job the tool will fit, because that knowledge exists only where the work is actually done.

What each group should do differently

Returning to the same three groups, in the same order:

  • Executives should change what they ask for before changing what they spend. Reporting that counts seats, licences and completed pilots describes procurement, not outcomes. Asking instead what people attempted, what they abandoned and what they kept produces a far more honest picture, and it is available without new systems.
  • Managers should distinguish protected time from mandated use. Gallup’s 2026 workforce data, drawn from 22,573 US employees, found that employees whose managers actively supported AI use were far more likely to report meaningful gains, outperforming individual usage frequency as a predictor. But “support” is doing heavy lifting in that finding. There is a substantial difference between protecting someone’s time to go looking and requiring them to route routine email through a model. The second is micromanagement in new clothing, and it will register as support on a survey.
  • Individual contributors should treat the fourth attempt as the experiment. The most common failure is a single unsatisfying result followed by a conclusion. Gallup’s data suggest that breadth of application, rather than frequency, tracks most closely with reported benefit: among employees using AI for one or two categories of task, 45 percent reported improved productivity, rising to 90 percent among those using it across seven or more.

We would add a caution to that last figure, as Gallup does. Employees who find value are also likely to look for more of it, so causation is unsettled. And roughly a quarter of employees who do not use AI report that they tried it and it did not help. Discovery is real, and it does not always find something.

The organizations pulling ahead are not the ones that bought earlier or spent more. They are the ones whose people accumulated enough attempts to find something, in workplaces where those attempts were treated as work rather than as slack.

That is a harder thing to put in a business case than a license count. It is also, on the current evidence, the variable that matters.

As discussed earlier, Steve Jobs was not arguing that customers should never be consulted. He said plainly that Apple researched its installed base and watched its industry. He was arguing that for something genuinely new, asking is the wrong verb, and showing has to come first. Three decades later, inside organizations rather than markets, the point holds with one modification. Nobody can do the showing on a knowledge worker’s behalf, because the value lives in the specific texture of their workday.

So the question worth putting to a team is not what they would use AI for. It is what they tried last week.

By Published On: August 11th, 2026Categories: AI Strategy, Future of workComments Off on Why can’t your AI pilot see the value your people are finding?

Share This Story, Choose Your Platform!

About the author : Charles

Charles Costa, MLIS is a researcher, strategist, and founder of Lexora Labs, where he works on AI adoption, knowledge management, and the future of expert