
What AI absorbs, what’s left for you
In early July, Commonwealth Bank of Australia and IBM both walked back AI-driven cuts they’d made earlier this year. CBA called its call-center reductions an “error.” IBM, after replacing much of its internal HR function with AI that could resolve only 94% of requests on its own, announced plans to triple entry-level hiring. Neither company reversed course because the AI was inaccurate in any obvious sense. They reversed course because it was confidently wrong in the ways only someone who’d actually done the job would have caught.
That gap, between what a model can calculate and what a person who’s been in the room knows, is becoming the central design question in AI deployment. Not whether to automate, but where the line actually falls.
The new hire who never sat in on a client call
AI is, in a sense, the new hire who nailed the interview but never sat in on a client call. It reasons from what’s written down, not from the tribal knowledge a team never got around to writing down, so it fails exactly where nuance matters most, and it fails confidently. That combination, wrong and confident, is worse than either problem alone, because confidence is what gets a bad answer shipped without a second look.
Charles Costa and Souvick Ghosh’s iConference 2025 research on generative AI in customer service found the same pattern holds industry-wide. Companies still need humans for effective service delivery because there’s no substitute for human judgment in mission-critical settings. The ones getting real performance gains aren’t the ones replacing agents with AI, they’re the ones using AI to clear routine volume so human agents can focus on the complex, judgment-heavy cases AI keeps getting confidently wrong.
The failures are architectural, not just a judgment problem
It’s tempting to treat this as purely a training or oversight issue, add a review step and the confident-wrong problem goes away. CIO’s recent analysis of enterprise AI reliability makes the case that most large-scale AI failures trace back to architecture, not model quality. Systems buckle under load. Prediction quality erodes as models drift over time. Even well-designed systems need a containment plan for the failures that inevitably slip through.
There’s a distinction worth sitting with here: accuracy isn’t the same thing as relevance. A model can correctly interpret a product requirements document written in July and still be dead wrong in August, once the requirements have moved on and nobody told the model. The fix isn’t a smarter model. It’s training the humans downstream to recognize when an answer has gone stale, and favoring narrowly-tuned, specialized models over one system trying to do everything.
What the research says about adding a human to check it
If the system itself is a source of risk, the obvious instinct is to put a person in the loop. The data on whether that actually helps is messier than it sounds. A 2024 systematic review and meta-analysis in Nature Human Behaviour, covering 106 experiments, found that human-AI combinations typically underperform the stronger of the two working alone. The exception was creative and content-generation tasks, where the combination reliably outperformed either one solo. Decision-making tasks showed the opposite: real losses, driven largely by over-reliance on the AI’s output.
Jakob Nielsen’s recent analysis for UX Tigers pushes the same point further. On convergent tasks, the kind with a calculable right answer, like forecasting or content judgment, a human reviewer is often working from partial context and gut instinct while the AI is running a consistent, structured methodology. The human override does more harm than good more often than most teams assume. The dividing line isn’t human versus AI. It’s convergent versus creative. Before building a human review step into any AI workflow, the question worth asking is whether the task has a right answer a model can calculate precisely, or whether it’s open-ended enough that judgment adds something the model genuinely can’t supply.
The labor market is already pricing this in
None of this is abstract for hiring. PwC’s 2026 Global AI Jobs Barometer found that the most AI-exposed junior roles are now seven times more likely than the least-exposed ones to require traditionally senior skills, leadership chief among them. That’s not a coincidence. As AI absorbs the routine, well-defined parts of entry-level work, what’s left for the person in the role is precisely the judgment and coordination that used to take years to earn.
The same data undercuts a common assumption: that AI-heavy companies are shedding headcount fastest. They’re not. The companies posting the biggest AI productivity gains are growing headcount and wages faster than their less AI-exposed peers, not because they’re optimizing for cost cuts, but because they’re using AI to unlock growth and new markets, with employees sharing in the upside.
Where this leaves founders
Every thread here points to the same fault line: not human versus AI, but convergent versus creative, routine versus judgment-heavy. None of it argues against automation. It argues against automating without first asking which side of that line a given task actually falls on. The organizations getting this right aren’t cutting headcount and hoping the model catches everything. They’re deciding, deliberately, where a person still has to be in the loop, and building compensation and career paths around exactly that judgment. For a founder planning headcount into next year, that’s the exercise worth running before the next AI rollout: does this task have a right answer a model can calculate, or does it need someone who’s actually been in the room before?
About the author : Charles

Charles Costa, MLIS is a researcher, strategist, and founder of Lexora Labs, where he works on AI adoption, knowledge management, and the future of expert






