Table of Contents
The best AI models in the world can win a math olympiad and still can't read a clock. For operations leaders, this nuance in AI capability is possibly the biggest challenge to build around in 2026 and beyond.
Last year, Gemini Deep Think earned a gold medal at the International Mathematical Olympiad. The same generation of frontier models reads an analog clock correctly about half the time – 50.1%, according to Stanford's 2026 AI Index.
I keep coming back to that pair of facts, because it explains almost everything about why AI initiatives in operations either compound or quietly die. Researchers call this the jagged technological frontier: capability doesn't arrive as a smooth rising tide. It arrives spiky. A model will do something genuinely astonishing and then fail at the adjacent task that looks, to a human, obviously easier.
If capability were smooth, adopting AI in operations would be a procurement decision. You'd wait for the number to get high enough and then buy. Because it's jagged, it's a design decision; and somebody has to know where the edges are.
This is the part I think most ops leaders are getting handed backwards. Almost every conversation I'm in starts as a capability question: can AI do this step? The actual question on the table is a placement question: where in this process does human judgment have to sit, and how do I make sure it stays there on the busiest day of the quarter?
The new ops stack, in my opinion, isn't a software category. It's a reallocation of who decides what.
The most expensive number in enterprise AI
You've probably seen the figure already: MIT's NANDA initiative reported that 95% of enterprise gen AI pilots produced no measurable P&L impact, against $30-40 billion spent. It gets quoted as evidence that the models aren't ready.
I don't think that's what it says at all.
McKinsey's late-2025 State of AI survey found 88% of organizations using AI in at least one function, and roughly 6% capturing 5% or more of EBIT impact from it. The variable separating those groups isn't model choice. High performers were about 3x more likely to have fundamentally redesigned their workflows. Only around 21% of gen AI users had redesigned any workflow at all.
So roughly 4 out of 5 companies bought a jet engine and bolted it to a horse-drawn cart, then expressed disappointment about the cart.
Stanford's Digital Economy Lab, in its Enterprise AI Playbook, goes into a study of 51 deployments that actually worked. In 77% of them, the hardest challenges were invisible and intangible: change management, data quality, process redesign. Not compute, or model quality. And my favorite finding in the entire report: 61% of successful projects included at least one prior failure whose costs never appear in the final ROI calculation.
Which means most of the ROI numbers you've read in a case study have been, in the gentlest possible sense, laundered. The failure was real, but they didn’t think to add it to the slide.
Klarna wasn't an AI failure, but a human one
Klarna is the story everyone cites, usually badly.
In 2024, the buy-now, pay-later financing company’s AI customer service chatbot supposedly did the work of roughly 700 human agents, cutting average resolution time from about 11 minutes to under two, and helped drop headcount by around 22%.
In 2025 the company reversed course and started bringing the humans back. In CEO Sebastian Siemiatkowski's own words to Bloomberg: “From a brand perspective, a company perspective, I just think it’s so critical that you are clear to your customer that there will always be a human if you want.”
Read that carefully, because it isn't a confession that the AI didn't work. What Klarna got wrong was which conversations needed a person: a placement error, not a capability error.
Commonwealth Bank of Australia made the same category of mistake at smaller scale, reversing a plan to cut 45 call-centre staff after its voice bot underdelivered and call volumes went up.
Now hold that against IBM's AskHR, which reached 94% containment and contributed to a 40% reduction in HR operating costs over four years. The difference is that IBM consistently described what it was doing as redesigning a service model, not adding a chatbot to one. Verizon did a version of the same thing and pointed the freed capacity at selling rather than at savings.
One set of teams asked what they could remove. The other asked where people should be. Those questions produce different companies and outcomes.
We're missing the vocabulary for designing around AI
Here's the most useful thing I've read all year, and it's almost boringly practical. Stanford's playbook found that successful deployments cluster into three oversight designs: collaboration, approval, and escalation.
Escalation – AI resolves, humans handle exceptions – dominates high-volume work where mistakes are recoverable.
Approval – where a human signs before anything moves – persists in claims, regulated finance, and healthcare, and will keep persisting, because a certified human is legally required to hold the liability.
Collaboration – humans and AI iterating together – holds in coding and creative work where quality is contextual and the loop is fast.
These are org designs. And the payoff for getting this right is not small.
In Stanford's sample, agentic implementations showed a 71% median productivity gain versus 40% for high-automation approaches, and only about 20% of implementations were fully agentic. Autonomy pays when human oversight is designed into the process, not when it's removed. Those look identical on a slide and behave nothing alike in production.
To slightly overshare, I've spent an unreasonable share of my life watching BTS choreography videos, and the thing that always amazes me is that the hardest part is never the individual moves but the formation changes: who slides to center on which count, who covers the gap left behind, who has to gracefully leap from the very end of one side of the formation to the other.
I think of operations in the same way. Almost nobody's process breaks because a single step was too difficult. It breaks at the handoffs.
A gate is not a suggestion in AI automation
McKinsey buried the real constraint in a single line: the scale of agentic adoption will be capped by how much oversight capacity humans can provide, "making governance itself a potential bottleneck to productivity."
I'd go further. Most human-in-the-loop AI, as actually implemented, is a notification. If your reviewer clears 400 approvals a day by clicking, you don't have oversight but mere ceremony attached to an audit trail that will one day be read aloud to you by a regulator.
The clearest illustration of this is Air Canada, whose chatbot invented a bereavement refund policy. When the customer sued, the airline argued that the chatbot was a separate legal entity responsible for its own actions. The British Columbia Civil Resolution Tribunal called that claim a "remarkable submission," since the chatbot was still part of their website, and awarded the customer damages amounting to C$812.02.
Victor Frankenstein's actual sin, in Mary Shelley's novel, was never building the creature. It was what he did the moment it opened its eyes: he left the room. Two hundred years later, a company stood in front of a tribunal and made the same argument he did; that thing acts on its own now. It didn't work then either.
Documented AI incidents rose to 362 in 2025 from 233 the year before. Without provenance, human-in-the-loop becomes human-in-the-blame-loop. A gate that can be bypassed under deadline pressure was never a gate but a preference.
Operations inherited the org chart instead of designing it
This is where the restructuring gets real, and where I think ops leaders are being handed authority nobody formally gave them.
GitLab tied its 2026 restructuring explicitly to the agentic era: up to three management layers removed in some functions, roughly 60 smaller end-to-end teams, approvals and handoffs rewired around agents. Shopify made AI a headcount gate: prove AI can't do the work before asking for people. Microsoft's research suggests leaders are starting to optimize something like a human-agent ratio rather than a traditional span of control. McKinsey describes teams of two to five people supervising 50 to 100 specialized agents, with humans "above the loop."
Those are org-design decisions. Layers, spans, career ladders, who learns what. They used to be made by the CEO and the CHRO. They're increasingly being made inside operations reviews, by people whose mandate says nothing about workforce design.
Which brings me to the cost I'd watch most closely, because it's the one that won't show up in any dashboard for five years.
The displacement panic might be overstated: Stanford found headcount reduction was the largest outcome in 45% of deployments, meaning the majority centered on hiring avoidance, redeployment, or no reduction at all.
But hiring avoidance has a victim, and it's the junior role. PwC's 2026 barometer found AI-exposed entry-level jobs are seven times more likely to demand traditionally senior skills. Employment for software developers aged 22 to 25 is down nearly 20% from 2024.
And here's the part that genuinely bothers me. In their paper “Generative AI at Work,” Brynjolfsson, Li, and Raymond found that AI assistance raised customer-support productivity 14-15% on average but up to 34% for novices, with minimal gains for experts. The technology is most useful to exactly the people companies are now declining to hire. We built a machine that compresses the learning curve, and are using it to remove the bottom of the ladder.
There's a counter-note I love, though, so I'll end on it. McKinsey found that employees without technical backgrounds learn to build agentic workflows about as fast as trained engineers, and specifically cited a French literature graduate. As someone who reads considerably more fiction than code: that is the most hopeful sentence in this entire body of research. The scarce skill turns out to be reasoning clearly about a process, not writing the automation. That skill is distributed far more widely than the current job market believes.
The question I'd actually ask when redesigning ops
At Moxo, we’re close to this human-in-the-loop problem. We build software around it: structured Human + AI processes that run across teams, clients, and vendors, with AI handling the automatable middle and humans holding gates that are required rather than optional.
And as someone close to the problem, here’s what I’d actually ask if I were sitting in an ops review tomorrow: Not what the AI can do. I'd ask three questions about every step we're automating: what happens when this is wrong, who finds out, and how fast we can undo it. If a process has good answers to those, you can be aggressive.
If it doesn't, no amount of model capability will save you. It'll just help you be wrong at scale, faster, with a cleaner audit log.
The jagged frontier isn't a phase we're passing through on the way to smooth. It's the working condition. Winning a math olympiad and knowing what time it is turn out to be different skills. So do executing a process and understanding one.
I’ll leave you with this: The operations leaders who come out of this decade ahead won't be the ones who automated the most. They'll be the ones who knew where to stand within the automations.

