
2026 Is the Year of Token Anxiety
I have noticed a new ritual appearing before people give an AI agent a serious task.
First, they check the model. Then the effort level. Then the speed. Then the remaining usage. They look at the reset time, wonder how much capacity the task might consume, and make a small calculation: should I start now, choose a cheaper model, or save the good tokens for something more important?
The work has not even begun, but the anxiety has.
This feels increasingly familiar in 2026. As agentic AI becomes part of everyday work, tokens are no longer an invisible technical detail. They have become something we can feel ourselves consuming.
2026 is the year of token anxiety.
Range anxiety, but for intelligence
Electric-vehicle drivers know the feeling of range anxiety: the fear that the battery will not carry them to the destination, even when the dashboard says there should be enough.
The problem is not simply battery capacity. It is uncertainty. The estimate changes with speed, temperature, terrain, driving style, and traffic. A new owner does not yet know what to trust. Thirty percent may feel comfortable on a familiar commute and dangerously low on an unfamiliar road.
Token anxiety has the same structure.
You may have usage remaining, but will it be enough to finish the task? Will a higher effort level solve the problem, or merely consume more capacity? Is the fast option economical, or does speed come from burning through the allowance more quickly? If the agent encounters a difficult test failure halfway through, will there be enough left to recover?
The important question is not:
How many tokens do I have?
It is:
Do I have enough intelligence to reach the destination?
The fuel gauge has entered the workplace
For most users, tokens used to be hidden. Developers paid for API calls, engineers optimized prompts, and finance teams received the bill. The person asking the question rarely thought about the machinery underneath.
Agents changed that relationship because they do more than produce a short answer. They inspect files, search documentation, operate tools, run tests, revise their work, and sometimes coordinate other agents. A single assignment can remain active for minutes or hours and consume an amount of capacity that is difficult to predict in advance.
At the same time, the controls have multiplied. We choose among models, reasoning or effort levels, standard and faster execution, local and cloud work, one agent or several. Each option promises a different combination of intelligence, latency, and consumption.
These controls are useful. They also transfer a new optimization problem to the user.
Before delegating the actual task, we are asked to become miniature capacity planners.
We do not know what a token buys
A litre of fuel and a kilowatt-hour are physical units. They are imperfect predictors of range, but drivers can gradually learn what they mean.
A token is harder to interpret. It is a unit of text and computation, not a unit of completed work. Ten thousand tokens might produce an excellent analysis, a failed attempt, a long conversation about requirements, or the final steps of a software feature that has already required much more.
The same apparent task can also consume very different amounts depending on the repository, the tools available, the number of mistakes encountered, and the standard of verification. “Fix the bug” might mean changing one line. It might mean spending an hour discovering that the bug is in a completely different system.
This makes the usage meter psychologically powerful but operationally weak. It tells us how much fuel is disappearing without telling us whether we are getting closer to the destination.
Every choice feels like it could be the wrong one
Model choice creates another layer of anxiety.
Select a smaller or lower-effort model and you may save capacity, but perhaps it will misunderstand the task and require several attempts. Select the most capable model and you may wonder whether you have used premium intelligence on work that a cheaper option could have handled. Choose speed and you may worry about cost. Choose efficiency and you may spend your own time waiting.
There is rarely enough information to calculate the perfect answer before the task begins.
This produces a strange form of decision regret. If the result is poor, we blame the model choice. If usage drops quickly, we blame the effort level. If the agent succeeds easily, we wonder whether we overpaid.
The optimal setting becomes obvious only after it is no longer useful.
The psychology of the reset
Subscription limits add their own behaviours.
When a reset is approaching, unused capacity can feel wasted. When the next reset is far away, the same capacity feels precious. A user with twenty percent remaining may postpone valuable work because they are afraid of needing those tokens later. Another may launch unnecessary tasks shortly before a reset because leaving capacity unused feels like losing something they already paid for.
This is not rational resource allocation. It is the familiar psychology of scarcity, expiration, and sunk cost applied to machine intelligence.
Different AI products package access differently, which makes intuition even harder to transfer. A habit learned in Codex may not apply to Claude Code. A personal subscription may behave differently from a metered API or an enterprise account. The user must learn a new fuel gauge for every vehicle.
When the company pays, anxiety becomes political
Token anxiety becomes more serious when the usage belongs to an employer.
If a company pays per token, employees may not know what responsible use looks like. Is a long agent run evidence of productive automation or careless spending? Is experimentation encouraged? Is there a budget per person, per team, or per outcome? Will someone receive a report showing who consumed the most?
Without clear norms, people invent their own. They choose weaker models, avoid exploratory work, shorten prompts, stop agents before verification, or quietly use personal accounts. A company can purchase powerful AI and still teach its employees not to use it through the ambiguity surrounding the bill.
The opposite is also possible. If usage feels unlimited and nobody connects it to outcomes, teams can generate enormous amounts of activity without creating meaningful value.
The answer is not tighter anxiety. It is clearer economics.
The optimization paradox
Some optimization is sensible. Easy tasks do not always need the most capable model. Repeated workflows should be made more efficient. Failed loops should not be allowed to consume resources forever.
But there is a point where token optimization becomes self-defeating.
If a knowledge worker spends fifteen minutes comparing models to save a few cents, the organization has optimized the cheaper resource. If an agent stops before running the tests because verification costs tokens, the apparently efficient run may create expensive human rework. If fear of using a reset delays a valuable task, unused capacity has become more important than the outcome it was purchased to create.
We can spend more human attention worrying about tokens than the tokens are worth.
This is why I believe the most useful metric is not consumption alone. As I argued in The Megatoken Economy, the better question is what economic value those tokens produce. A costly run that completes a valuable workflow can be efficient. A cheap run that produces nothing useful is not.
From a fuel gauge to a confidence gauge
Better AI products should not merely show a percentage and a reset clock. They should help users answer the question that actually creates anxiety: can I finish this task?
That might mean estimating the likely capacity required, warning before a selected mode is disproportionate, preserving enough headroom for verification, or allowing a task to continue gracefully with a different model. It might mean explaining whether “remaining usage” is a hard limit, a rolling allowance, or an estimate.
The interface should communicate confidence, not just inventory.
Organizations have work to do as well:
- Provide sensible defaults. Most employees should not have to solve a model-routing problem before every prompt.
- Match effort to stakes. Use more capable settings when errors are expensive, the task is hard to verify, or the result has high value.
- Define experimentation budgets. People should know how much freedom they have to learn without fearing an invisible cost review.
- Measure outcomes alongside usage. Cost becomes meaningful only when compared with time saved, quality improved, risk reduced, or value created.
- Design for recovery. Running low should not mean abandoning a half-completed task and losing all the context invested in it.
Individuals can adopt one simple rule: choose a reasonable default, change it when the stakes justify the change, and evaluate the result rather than obsessing over the meter.
Anxiety is a sign of transition
Early electric-vehicle owners had to think constantly about chargers, routes, temperature, and remaining range. As batteries improved, charging networks expanded, and estimates became more trustworthy, much of that mental burden declined.
AI will probably follow a similar path. Capacity will become cheaper. Routing will become more automatic. Products will learn which model and effort level a task requires. Organizations will develop clearer budgets and expectations.
But in 2026, we are still new owners. The gauges are unfamiliar, the estimates are uncertain, and the charging networks all have different rules.
That is why token anxiety matters. It is not a complaint from people who want unlimited AI. It is evidence that agentic AI has become real enough to depend on, while its economics and interfaces remain difficult to trust.
The industry has made intelligence visible as a consumable resource. Its next challenge is to make that resource understandable.
Until then, we will keep looking at the meter before beginning the journey.
Not because we are afraid to use AI, but because we are no longer sure whether we have enough AI to arrive.





