800.553.8359 info@const-ins.com

Good morning. Intelligence is getting radically cheaper, but enterprise AI bills keep climbing anyway.

That was the central paradox McKinsey senior partners Tanguy Catlin and Lari Hämäläinen tackled on Tuesday during a McKinsey Live virtual session, “Improving the Economics of Agentic AI,” which highlighted the firm’s State of AI in 2026 survey.

“Intelligence at a certain capability level is getting a lot more affordable,” said Hämäläinen, who is also a leader in McKinsey Digital.

For example, GPT-4 launched in early 2023 at $60 per million output tokens. Today, models with roughly GPT-4-class performance on established benchmarks can be served at a fraction of that cost—in some cases hundreds of times cheaper, with prices falling from tens of dollars per million output tokens to well below a dollar, he explained.

But as the cost of producing a unit of intelligence collapses, the amount businesses consume is exploding.

Models are getting cheaper per unit of capability while enterprises ask them to perform vastly more reasoning and work, especially through autonomous agents. AI vendors are also capturing some of those efficiency gains through higher margins, Hämäläinen said.

In software development, for example, AI agents can repeatedly inspect, modify and rewrite entire codebases, generating far more code than a human developer would typically touch, he said.

The economics of agentic AI

Companies are only beginning to understand the economics of agentic AI, Hämäläinen said. Unlike traditional software, where the cost of running a task is relatively predictable, agents can take different paths to the same result, making costs highly variable. The same task can cost up to 30 times more from one run to another, he said.

Much of that cost comes from the reasoning and repeated refinement behind the final output. System design matters: choices such as using a single agent versus multiple agents can dramatically affect the cost of completing a task, Hämäläinen said.

AI cost management has become more about engineering systems that don’t waste tokens in the first place.

Rather than measuring agents by cost per token, he said leaders should evaluate them at the task level: how much a task costs to execute, how often the agent succeeds, and how much human time is required to verify its work.

As a rule of thumb, an agent can make sense when the time required to verify its output is a small fraction of the time it would take a human to complete the task from scratch, he said. If a task takes a human an hour, but an agent’s work can be verified in six minutes, an agent with a success rate above 10% could already begin to create value, he added.

The bigger challenge, he said, is redesigning the surrounding workflow to take advantage of the capacity the agent frees up.

Meanwhile, “The truth is, there is no single cost lever,” said Catlin, who is also a director of the McKinsey Global Institute. He identified three areas where companies can manage AI spending:

—First, companies need visibility into which use cases, business units, agents, models and users are driving spend.

—Second, they need to optimize workflows by matching model complexity to the task, routing requests to appropriate models, caching reusable context and limiting unnecessary tool calls and agent loops.

—Third, they need greater sourcing discipline, including removing unused licenses, managing quotas, negotiating provider terms and avoiding excessive dependence on a single model or vendor.

But Catlin cautioned that the goal should not be indiscriminate cost-cutting. Companies should determine where AI spending delivers the highest returns and optimize toward those returns.

Sheryl Estrada
Sheryl.Estrada@fortune.com

This story was originally featured on Fortune.com