Essay
The Year Inference Got Cheap
Quietly, the most important number in AI has been falling off a cliff. The cost of a unit of machine reasoning, a token, has dropped by more than an order of magnitude in a short span, and it is still falling. Reasoning models that think before they answer arrived at the same time. Thinking, in other words, got both better and cheaper at once. Most enterprise AI strategies were written when it was still expensive, and they have quietly gone out of date.
Why this changes the plan
When intelligence is expensive, you ration it. You use AI only on the few high-value calls, you keep prompts short, you avoid letting a model loop. Every one of those instincts becomes wrong when the price collapses.
Cheap inference means you can afford to let a system think longer, check its own work, try several approaches, and call a model many times inside a single task. The behaviours that were too expensive last year, verification, self-correction, multi-step reasoning, are now just line items. The question shifts from "can we afford to use AI here" to "why are we not using it everywhere it removes toil."
The trap on the other side
Cheap does not mean free, and cheap per call does not mean cheap at scale. When a single user action quietly triggers fifty model calls, your bill can grow faster than your value. Falling prices reward volume, and volume without discipline is how AI programs blow their budget while looking productive.
So the discipline is not rationing anymore. It is instrumentation. Know your cost per outcome, not your cost per token, and watch it like a unit economic.
What to do now
- Revisit the use cases you killed for being too expensive. Many are now viable. Your no-list from last year is your opportunity list this year.
- Design for thinking, not just answering. Let systems verify and retry. The extra calls are cheap, the wrong answer is not.
- Track cost per outcome. Falling unit prices hide rising totals. Measure the thing the business actually pays for.
The teams that win are not the ones who spent the least on tokens. They are the ones who noticed the price moved, and rebuilt the plan around it.
So before your next AI roadmap: which of your assumptions were priced at last year's cost of thinking?
