Most Tasks Don't Need Frontier Models
Most tasks don't need the smartest frontier model available. The real bottleneck isn't intelligence, it's how much of your own attention you have left to check the work. On the ...
Most tasks don’t need frontier models
Every task has an intelligence threshold, a level of reasoning below which the output is wrong or unusable, and above which you're paying for reasoning the task doesn't need.
Drafting a standard client email, cleaning up a spreadsheet, summarising a meeting, these have low floors.
The economically correct model for any given job is therefore the cheapest one that clears its floor, and most tasks have low floors.
The most expensive frontier intelligence is often being wasted on tasks that could be decomposed, routed, constrained or skipped.
When the token subsidies end
You're not really paying what your AI usage costs to produce, because the providers are selling inference below cost, funded by venture capital.
That's why everyone routes everything to the frontier model. When intelligence is priced at almost nothing, there's very little incentive to match a task to a cheaper model.
But when the frontier model subsidy ends and access is metered, every user suddenly sees a cost attached to each task.
Amdahl's law says the speed of a whole system is set by its slowest necessary step, not its fastest. Most workflows have a small number of genuinely hard steps and a large body of easy ones, so the leverage in cost comes from spotting the few steps that justify the expensive model and routing everything else down.
Attention Management
There's a well-known limit on working memory, the number of distinct things a person can actively supervise at once, and for most people that number is close to two.
Push past that limit and we don't do three things at sixty percent each. We do all of them poorly, because we're not really running them in parallel, we're switching between them fast and paying a small tax on every switch.
Neither more money nor a better model raises that limit. It's a property of attention itself, not of the tools you're using.
Producing output with AI agents has gotten cheaper and faster, but your judgment didn't. Adding more agents past the point where you can actually judge them just adds queue depth, not necessarily a stream of higher quality completed work. The ceiling is set by how fast you can render judgment, and that didn't move as producing output became cheaper.
Sophie Leroy's 2009 finding on attention residue sharpens what that tax looks like. Switching away from a task leaves part of your attention behind on it, a film that coats the next judgment. Supervising a spread of agent loops means living in permanent residue, where every glance at one loop degrades the quality of attention available to the next.
That means attention is the genuinely non-fungible input of the agentic era. Every other input scales with money. There's no market where you buy back your own attention, so it's the true unit of cost.
Loops
Frontier intelligence is artificially cheap for now and about to get a price, attention is permanently scarce and isn't priced by any market because it can't be bought, and a loop is the one object that manages both.
It’s the design of repeatable human and agent workflows where context, action, review and feedback are deliberately structured.
It rations your attention by converting a continuous supervisory demand into a few checkpoints selected in advance, so instead of watching constantly, you check in at a small number of points you chose ahead of time. And it rations intelligence by letting each step carry the cheapest model that clears its floor.
That dual design sits in process and systems thinkers, who already know where a workflow's hard steps are and where a human needs to step in.
So, a model gives you capability, a loop gives you control.
Most people are still treating AI work as a conversation. They ask, inspect, correct, ask again. That works when the task is small, but it collapses when there are many tasks, many agents, many outputs, many models and many possible next steps.
The loop changes the work from prompting to operating.
So, the deeper idea is that AI maturity has three stages.
First, people chase better prompts.
Second, people chase better models.
Third, people build better loops.
That third stage is where the business value likely appears, because it converts AI from an impressive interaction into an operating system for repeatable work.
What this all means for jobs and growth?
A lot of AI doomerism is based on the idea that if each worker now does the work of ten, you therefore need a tenth of the workers.
But I don’t think it’s quite that simple.
What remains scarce is judgment, the human looking at the output and deciding if it's good. Producing output got cheaper and faster, so the value of the work stops being gated by how much you can produce, but is gated by how much judgment you can bring to bear.
Judgment lives in people, so if you want more verified output you need more people, each one a fixed-size bucket of the one thing that's now scarce, another two threads of supervision capacity, another command post over a fleet of agents.
We’ve seen Jeavons paradox play out like this before.
ATMs were supposed to wipe out bank tellers, but teller numbers went up for decades after ATMs spread.
What happened is each branch needed fewer tellers, so branches got cheaper to run, so banks opened a lot more branches, and the teller's job shifted off cash-counting and onto the stuff a machine couldn't do, sales and relationships and handling the edge cases.
But, arming every employee with a swarm of agents is only affordable if those agents run on right-sized models, if every token is priced at the frontier instead, giving fifty people a fleet each is ruinous and the hire-more theory dies on cost.
So, cheap-enough intelligence per task is what lets a firm afford to put generation capacity behind every judgment-holder. And, attention caps out per person, so once you've maxed the agents one human can actually supervise, the logical lever left to grow output is another human.
The big events to watch out for won’t necessarily be the model provider IPO’s, but the end of the token subsidy which is sure to follow, because that’s what forces every one of us to ask, for the first time, how much intelligence each task actually needs.
We got into this and more this week in the Big Episode 62 of The Good Stuff.