
The hidden cost of verifying AI outputs
AI is supposed to save time. The part we rarely measure is how much of that time gets spent making sure the answer deserves to be trusted.
A 2025 study of 319 knowledge workers offers one uncomfortable number. Researchers asked participants how they used generative AI across 936 real work tasks. One hundred fourteen said they cross-checked AI output against an outside source. Only 23 checked the sources the AI itself cited.
It would be easy to read that as carelessness. Give people a tool that produces confident answers instantly, and eventually they stop doing the work required to verify them. But I think the research points somewhere more interesting.
People are adapting rationally to a technology that makes producing information dramatically cheaper while leaving the cost of evaluating that information largely unchanged. AI can create an answer in seconds. Deciding whether that answer is good can still take ten minutes, an hour, or require someone with years of experience. That gap is becoming one of the hidden costs of AI adoption.
The eleven hours saved come with an asterisk
The Work AI Index 2026 surveyed 6,000 digital workers across the United States, United Kingdom, and Australia. Workers reported that AI automation saves them roughly eleven hours every week. That is the kind of number that makes its way into an executive presentation.
Another number from the same research probably does not. Those workers also reported spending an average of 6.4 hours a week on what Glean calls “botsitting”: supplying missing context, checking outputs, debugging mistakes, rerunning prompts, and cleaning up AI-generated work. That represents 37% of the time workers spend interacting with AI, slightly more than the share spent actually using AI to produce work.
The biggest component is context. That tracks with what anyone doing substantive work with these systems eventually discovers. Ask an AI tool to map a database schema, analyze a market, interpret internal research, or find the right dataset, and the first answer is often not obviously wrong. It is simply missing everything about your particular situation that was never available to the model.
The system does not know why your team stopped trusting one dataset three months ago, which edge case broke the last implementation, or that a requirement changed in a meeting and never made it into the documentation. Getting the model to understand enough context to produce something useful is part of the work. Then you still have to check it.
This is where the productivity calculation starts getting slippery. The eleven hours saved are visible because they can be framed as work the AI completed. The 6.4 hours spent making the AI useful are scattered across prompts, corrections, source checks, comparisons, and judgment calls. One gets counted as productivity. The other disappears into the workday.
Fluency changes the decision to check
The verification problem gets harder because generative AI does not simply give people information. It changes how that information feels. In research from Microsoft Research and Carnegie Mellon, 319 knowledge workers described how they used generative AI across 936 real tasks. Twenty-three said they manually verified sources directly referenced by the AI, while 114 said they cross-referenced the information against outside sources.
The more consequential finding was what predicted critical thinking in the first place. Higher confidence in the AI was associated with less critical thinking, while higher confidence in one’s own ability was associated with more. That creates a peculiar problem for systems whose defining feature is their ability to sound competent.
A poor answer gives you reasons to investigate. It contradicts itself, misses something obvious, or sounds uncertain. A fluent wrong answer removes those warning signs. The better the response reads, the easier it becomes to accept the premise that there probably is not much left to check.
This is not necessarily laziness. Verification competes with every other demand on a person’s time. When an answer appears credible and the task feels low-risk, spending another fifteen minutes independently validating it can look irrational. The difficulty is that fluency is not evidence. It is presentation, and generative AI has become extraordinarily good at presentation.
Checking the source is not the same as checking the claim
Even opening the citation does not always solve the problem. An AI system can produce a misleading conclusion without inventing anything. It only needs to summarize a set of sources whose incentives, assumptions, or perspectives all point in the same direction.
Search for a minor medical problem, for example, and an AI summary may draw heavily from specialist practices. Those pages can be factually accurate while naturally emphasizing the situations in which professional treatment is appropriate. Thication is harder. You also have to ask whether the source gives you an independent reason to believe the conclusion.
In other words, provenance matters, but provenance alone is not judgment. AI can synthesize those sources faithfully and still leave someone with the impression that a relatively ordinary problem requires a specialist.
Nothing was hallucinated. A person can open every citation, confirm that the source is real, and still walk away with an answer that is skewed by the material selected. That matters because “check the source” sounds like a binary action. The source exists or it does not. The source supports the claim or it does not.
Putting a human in the loop does not guarantee human judgment
The standard answer to AI reliability problems is to keep a person in the loop. That helps only if the person continues to treat the decision as theirs.
A 2026 study published in PNAS Nexus tested this with 1,339 teachers. Participants evaluated student work alongside an intentionally incorrect grade. The work and the grade were identical; researchers changed only whether participants were told the recommendation came from a human or an algorithm.
For an unfairly harsh recommendation, the grading error became 22% larger when the same score was labeled as AI-generated. The people who deferred most strongly were also not the group most people would expect. Deference was particularly pronounced among younger, highly educated, technologically confident participants.
Greater comfort with technology did not necessarily produce greater resistance to it. The AI label altered the perceived competence and responsibility of the system, making the call feel a little less like the teacher’s to make.
That is a much harder governance problem than simply requiring review. A person can technically approve an AI-generated decision without exercising much independent judgment at all. The useful question is therefore not whether a human appears somewhere in the workflow, but how much of the decision that human still believes they own.
That boundary can move quietly. The system starts by drafting, then it recommends, then it ranks, then it evaluates. Each additional step feels like a small efficiency gain. Eventually assistance becomes deference without anyone formally deciding to transfer responsibility.
The reference layer can fail too
There is another assumption underneath verification: when something looks wrong, we can move upstream. Check the citation. Read the original. Consult the database. Find the precedent.
That works only while the reference layer remains trustworthy.
Damien Charlotin’s AI Hallucination Cases database tracks court decisions involving AI-generated hallucinated material, including fabricated case law, false quotations, and misrepresented authorities. The database now contains cases across dozens of jurisdictions and continues to expand.
The embarrassing filings and sanctions attract attention, but the more important risk is what happens after false information enters a system designed to function as a reference. Law runs on precedent. Research runs on prior literature. Businesses rely on policies, internal documentation, databases, reports, and knowledge bases. AI systems increasingly retrieve from the same material.
Once fabricated information enters one of those layers, verification can reproduce the error instead of catching it. The next person may do exactly what we tell responsible AI users to do: question the answer, find the source, and confirm that another document says the same thing. What they cannot see is that the error happened one layer earlier.
A hallucination inside a private chat is a contained mistake. A hallucination copied into a filing, report, knowledge base, policy document, or other durable record becomes part of the environment everyone else checks against. That turns AI accuracy from an individual productivity issue into a knowledge management problem.
Production scales faster than trust
Hiring shows what this looks like when the economics change at scale. Greenhouse analyzed more than 640 million applications for its 2026 recruiting benchmarks. In 2022, the average recruiter handled 146 applications annually. By 2025, that figure had reached 746. Over the same period, the average recruiting team fell from 10.43 recruiters to 4.62, while time to fill increased 37% anyway.
Greenhouse does not claim that AI alone caused those changes, and it would be wrong to make that leap for them. But the pattern illustrates the larger problem clearly. Technology made producing an application easier. It did not make evaluating an applicant easier at the same rate.
Once polished applications become cheap to create, polish stops carrying as much information. Recruiters have to look elsewhere for the signal. Does the candidate actually have the experience described? Was the application meaningfully tailored? Does the portfolio represent the person’s capabilities? Can they explain the work? The production cost falls, and the verification cost moves downstream.
The same dynamic will not stop at recruiting. AI makes reports cheaper to produce, proposals cheaper to produce, research summaries cheaper to produce, code cheaper to produce, and analysis cheaper to produce. But every recipient still has to answer the same question: should I believe this? Generative AI scales production much more easily than organizations can scale judgment.
Good, easy, trusted. Pick two.
Put these findings together and I do not see a story about everyone suddenly becoming careless. I see a series of reasonable adaptations whose costs land somewhere else.
A worker spends more time feeding the model context because doing so produces a better answer. Another skips a source check because the output looks credible and there is a deadline. An expert gives an algorithm slightly more authority because it is supposed to be better at the task. Someone verifies a citation without realizing that the reference layer has already been contaminated. A recruiter receives five times as many applications with half as many colleagues available to review them.
Each decision makes sense locally. The system becomes harder to trust collectively. That suggests a useful tension for AI-mediated work: good, easy, trusted. Pick two.
AI can make good-looking work dramatically easier to produce, but making production easier does not automatically make the result easier to trust. In some cases it does the opposite, because the effort previously required to create something was itself part of the signal. Once that signal disappears, the receiver has to compensate with more scrutiny. Critical thinking is therefore not just a soft skill employees should be reminded to use. It is operational capacity, and capacity has a cost.
Organizations need to budget for verification
“Trust but verify” remains good advice. What is missing is an operating model for the second half.
Every source someone opens, every assumption they test, every result they reproduce, every edge case they investigate, and every AI recommendation they independently challenge consumes time. If an organization expects that work to happen but does not allocate time for it, one of two things occurs: either employees absorb the cost invisibly, or they gradually reduce how much checking they do until the workload becomes manageable. Neither outcome should be surprising.
The answer is not to verify every AI output with equal intensity. That would erase much of the productivity benefit the technology can provide. A low-stakes rewrite does not deserve the same review process as a financial analysis, legal recommendation, hiring decision, medical answer, or executive briefing.
The answer is to decide where verification is worth paying for. That means defining which outputs require independent sources, which decisions require subject-matter review, where provenance needs to be preserved, when AI-generated material is allowed to enter an authoritative knowledge base, and who retains responsibility for the final call.
It also means being more honest about productivity. If AI saves an analyst four hours drafting a report and the analyst spends two hours checking the claims, correcting the context, and tracing the citations, the productivity gain is not four hours. That does not make the technology unsuccessful. It makes the accounting more accurate.
The organizations that get this right will not be the ones that eliminate verification in pursuit of maximum speed. They will be the ones that understand exactly which forms of verification create enough trust to justify their cost. Because somebody is already paying for the checking. Most organizations simply have not decided who. And that may be the real lesson behind twenty-three out of three hundred and nineteen.
Trust, but verify still works. We are only beginning to understand the price of the second half.
About the author : Charles

Charles Costa, MLIS is a researcher, strategist, and founder of Lexora Labs, where he works on AI adoption, knowledge management, and the future of expert






