Daniel Fallmann is founder and CEO of Mindbreeze, a leader in enterprise search, applied artificial intelligence and knowledge management.
For the better part of two years, enterprise AI conversations have centered on model performance, accuracy, security and workforce transformation. Those questions still matter. However, a different and more urgent challenge is emerging inside organizations that have moved past pilots and into real deployment at scale: the economics of AI consumption itself.
Token usage is quietly becoming one of the fastest-growing line items in enterprise technology budgets, and most organizations aren’t yet equipped to manage it.
From Predictable Licensing To Variable Consumption
Traditional enterprise software followed a familiar economic logic. Licensing costs were largely predictable, and per-user costs often declined as adoption grew. GenAI breaks that pattern.
Every prompt, retrieval, reasoning step and autonomous agent action consumes tokens that translate directly into cost. Individually, these costs look negligible. At enterprise scale, across thousands of employees and growing fleets of AI agents, they compound quickly and unpredictably.
Many organizations are now dealing with “AI sprawl,” in which employees adopt multiple overlapping AI tools with little centralized coordination. This produces duplicated work, fragmented workflows and escalating token consumption without a corresponding gain in productivity.
Agentic AI Changing The Cost Equation
The challenge intensifies as organizations move from simple assistants to agentic AI. Unlike a single chat interaction, an autonomous agent rarely makes one model call. It retrieves information, invokes external tools, reasons through intermediate steps, validates its own outputs and often repeats parts of that cycle before a task is complete.
Each of those steps consumes tokens. Multiply that across hundreds or thousands of concurrent workflows, and AI spending starts to behave less like software licensing and more like cloud infrastructure consumption: elastic, usage-driven and easy to lose control of.
Vendors are already responding. Axios reported in June 2026 that Databricks rolled out enterprise controls specifically designed to cap AI spending and monitor usage across providers, after some organizations found themselves with AI bills reaching tens of millions of dollars per month. The broader trend has been dubbed “token maxxing” to describe how aggressive, uncoordinated AI adoption can drive token consumption that delivers little measurable business value.
A Familiar Pattern, A Faster Curve
Enterprise technology leaders have seen this movie before. Cloud computing had made infrastructure trivially easy to provision, and many organizations expanded usage well beyond what their budgets or governance structures could support. FinOps emerged as a discipline precisely because consumption had outpaced oversight.
AI is now repeating that cycle, except it’s doing it faster because token consumption can scale invisibly the moment AI’s embedded into everyday workflows rather than provisioned through a visible request process.
Many European enterprises expanding AI deployments are already diversifying providers while placing a greater emphasis on cost management and infrastructure governance due to autonomous systems consuming far more tokens than originally projected.
Why Knowledge Architecture Matters To The Cost Equation
In my experience working with enterprises on knowledge management and applied AI, I’ve found that a significant share of unnecessary token consumption traces back to one root cause: AI systems that lack direct, governed access to the right information and instead compensate through repeated retrieval, broader context windows and redundant reasoning steps. When an AI system has to search, re-search and reconcile fragmented or duplicated knowledge across disconnected repositories, every one of those extra steps is a token cost the organization didn’t need to incur.
This is why I believe the next phase of enterprise AI governance must start with the information architecture underneath the AI. Organizations that connect their AI systems to a unified, well-governed knowledge foundation, rather than letting agents search blindly across silos, can naturally consume far fewer tokens to reach the same or better outcomes. Efficient retrieval is increasingly a cost discipline.
Measuring Value, Not Just Usage
Many organizations still measure AI success through adoption metrics: how many employees use AI, how many prompts are submitted, how many workflows have been touched by automation. While those metrics were useful during experimentation, they become misleading once AI spending reaches meaningful scale because consumption isn’t the same thing as value.
Rewarding usage alone can unintentionally encourage longer prompts, redundant model calls and AI invoked simply because activity has become a proxy for innovation, echoing the same mistake many organizations made when they equated more cloud infrastructure with digital transformation.
The more important question for executives is whether the organization’s AI usage is generating measurable business outcomes relative to the resources it consumes.
Token Governance As The Next CIO Imperative
AI oversight is expanding beyond security, compliance and privacy to include economics. Executives need visibility into which models are being used, how frequently agents invoke them, which workflows generate the greatest business value and where token consumption is happening without a return.
The opportunity in front of us remains substantial. Agentic AI can genuinely eliminate operational friction, accelerate knowledge work and improve decision-making across the enterprise. However, realizing that opportunity will depend less on maximizing model usage and more on maximizing business outcomes per token consumed, built on a foundation of governed, unified access to enterprise knowledge.
The organizations leading the next phase of enterprise AI will be the ones that operate AI with the greatest discipline, treat token governance with the same rigor they once applied to financial governance and build on knowledge infrastructure designed for efficiency from the start.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


