Tokenomics Is Now a Management Discipline

Ten takeaways from the collective intelligence of the Executive Technology Board

The metric changed, and most enterprises have not caught up

Cost per token is the wrong measure and is now widely recognized as such by the leaders furthest along. The measure that matters is cost per output, per workflow, per feature, or per business outcome. A large token bill is entirely rational if it lowers the total cost of delivering the work. A small token bill can still be pure waste if it produces nothing the business can count. Enterprises that are still negotiating unit price without a view of unit output are optimizing the wrong variable, and they will not discover that until the pricing cycle turns.

Attribution is the primitive that makes everything else possible

The most operationally serious enterprises I have seen tag every token at the point of issuance, allocate it to the business unit that owns the consuming user or system, and charge it through total cost of ownership. If a business unit wants to spend more on tokens, it finds an offset elsewhere in its operating cost. Daily consumption taps are set as soft limits rather than hard stops, because the objective is to surface the pattern rather than interrupt the work. Without attribution, every other cost control is theater, because nobody can be held accountable for a number they cannot see.

Run cost is migrating from the technology budget to the business P&L

The emerging funding model is consistent across sectors. Central technology funds experimentation, platforms, infrastructure, and development. Business units assume responsibility once a capability enters production. This does more than move money. It converts AI from innovation spend into normal operating economics, and it aligns incentives with business outcomes rather than with adoption metrics. The enterprises that have made this move report noticeably better discipline in what gets promoted to production in the first place.

Nobody optimizes during experimentation, and that is correct

The clearest shared operating rule in these conversations is that speed matters more than efficiency during experimentation, and the priority inverts the moment something reaches production. At that point engineering effort shifts deliberately toward reducing inference cost while holding quality constant. The reported reductions from prompt engineering, model selection, and workflow redesign after deployment are material. Optimizing too early kills good ideas in the business case phase, which several leaders described as the place where most of their ideas already die.

The expensive lock-in is at orchestration, not at the model

Considerable executive attention goes into model choice, and comparatively little into the layer above it. That allocation is backwards. Routing work to the right model at the right cost is where the economic discipline actually lives, and the pattern that works is composing multi-stage workflows so that the expensive model handles design and reasoning while smaller models take over once the workflow is understood. The harder point is that frameworks, accumulated context, embedded workflows, and skills are where switching cost accumulates. Models are increasingly swappable. Orchestration is not.

The costs that surprise people are second-order

Three surfaced repeatedly and almost nobody had modeled them in advance. One enterprise found that a security overlay doubled the token cost of the workflow it was protecting, which turns cyber as a percentage of AI run cost into a board question with no available benchmark. Poor software architecture raises token burn materially, because the model spends effort understanding and unwinding messy code, which makes clean architecture an economic decision rather than an engineering preference. Agent-to-agent traffic and audit retention add a third layer that few enterprises have priced at all.

The exposure is the price assumption, not the consumption

Unit costs are falling, in some cases sharply, and this is the fact most often used to dismiss the concern. It is the wrong fact to rely on. Volume is growing faster than unit cost is declining, providers are widely believed to be subsidizing on the path to public markets, and the next pricing cycle is expected to be materially less favorable. Token economics will also produce geographic pricing, local model demand, and sovereignty-driven cost differentials. The working position in most of these rooms is to run every business case at three times current prices, and to treat any deployment that fails that test as a deployment with a shelf life.

General-purpose adoption still has no honest metric

Traditional AI projects have business cases. General-purpose assistants do not, and this is the least resolved question in the entire discussion. Candidates under consideration include departmental productivity, hours saved, innovation funnel contribution, process improvement, client and employee retention, leadership adoption, and token efficiency. No consensus has emerged. What has emerged is agreement that login counts and license deployment are close to meaningless, and that any metric which rewards consumption will produce token maximizing as a badge of honor. The right question is whether AI has changed how the work gets done.

The winning posture is efficiency, not restriction

There is a real temptation to control cost by controlling access, and the enterprises making the fastest progress have refused it. Their approach is to coach heavy users, improve prompt quality, route to cheaper models where quality permits, and deploy specialized tools instead of general-purpose ones. The framing I heard that captures it best is that the objective is business leaders managing outcomes rather than IT policing employees. I have heard of at least one team told to stop using AI on token-cost grounds, and I would treat that as a failure of architecture rather than a model to follow.

This is a CFO conversation before it is a CIO conversation

A token is a unit of language, and enterprises have been building software with language for decades. Requirements argued in meetings, specifications written in documents, review comments in email. That language production has always been the real cost of software, and it has never once been metered, because it sits inside salary. Compute tokens are metered, itemized, and visible on an invoice, which is why the smaller cost feels alarming while the larger one stays invisible. Compare the total cost of the outcome under both models and compute usually looks cheap. The consequence is that architecture is a vehicle for an economic outcome, and the trade at the center of it, accepting a visible metered line in exchange for reducing an invisible one, is a trade only finance can book. The CFO needs to be in the architecture conversation, not briefed after it.

In summary

The enterprises that will absorb the next pricing cycle gracefully are the ones making architectural decisions in the next twelve months with the economics explicit. The ones that will not are treating this as an engineering question with a procurement answer. Tokenomics is not a cost-control exercise. It is the discipline that determines which AI deployments survive contact with real unit economics, and it belongs on the board agenda for exactly that reason.

Executive Technology Board (c)