On a recent Cisco earnings call CEO Chuck Robbins sounded like he was confessing to a problem rather than celebrating growth. Robbins said the company’s internal chatbot now sees daily use from about one third of staff and that "the token usage is getting pretty, pretty crazy," a line that captured how quickly AI consumption has moved from experiment to budget headache. Across tech teams, firms from Microsoft to Priceline have tightened access, revoked developer licenses, and imposed hard token limits after early 2025 usage exploded. The scramble matters because companies now face a blunt trade-off between buying more model capacity and keeping human roles intact.
In conference rooms, developer desktops and customer support queues the frontier models that promised productivity gains are now driving finance headaches.
From open access to hard limits
That shift has been sudden. Engineers and business teams folded Anthropic's Claude, OpenAI's newer models and Google's top-tier models into daily workflows for code generation, email drafting, feedback analysis and agentic automation. Those models, released late last year and this spring, brought capability improvements that tempted broad unmetered use. But higher capability came with higher per-token costs, and firms that let teams test freely found usage multiplying and bills following.
Some of the rollbacks are strikingly concrete. Microsoft revoked developer licenses to Claude Code months after enabling them. Priceline said routine vendor renewals spiked several times their prior cost. A single unnamed company reported a vendor bill that reached roughly $500 million after failing to set usage limits. And at Priceline a senior IT finance director compared the dynamic to an addiction, saying the tools "let you try it to get you hooked on it, and now you’re kind of beholden to it." Engineering teams at multiple firms have logged monthly token spends in the tens of thousands of dollars.
Not every organisation saw only rising costs. Communications company 8x8 estimated it saved about $5 million annually by cancelling subscriptions to software and education services that AI could replace, and its chief transformation officer said the company’s annualized Claude bill remains "well below" that figure. But other large enterprises painted a starkly different picture. Royal Bank of Canada’s CEO disclosed token consumption surged 500 percent over six months, and executives at analytics and cloud companies said top engineers can run up "thousands of dollars a month or more" in token spend.
The budgeting consequence has been visible on public calls and in filings. Roughly 300 companies raised token-related questions in April and May earnings calls, compared with 93 a year earlier, an indication that token budgeting has moved from a niche operational concern to a routine C-suite and board item.
Industry responses are clustering along three tracks: tooling, governance and standards. A market of startups and incumbents is selling cost-visibility platforms, model-routing services and controls that let finance and engineering pick the cheapest adequate model for a task. Internally, firms are implementing usage quotas, breaking out token costs by team and adding billing alerts. Procurement teams are negotiating narrower contracts and demanding audit trails.
Some executives are reintroducing human labour where AI costs don't yet justify replacement.
The Linux Foundation announced plans to stand up a Tokenomics Foundation to push a common vocabulary and guardrails for token accounting and efficiency, an effort framed as similar to how FinOps brought discipline to cloud spend. J.R. Storment, executive director of the FinOps Foundation, said companies told him in April and May they were multiples over their budgets and "we started hearing existential crises," a shift that moved discussions from model capability to cost auditability and control.
Enterprise AI vendors and builders also point to technical levers that can cut bills. One enterprise AI CEO said roughly 95 percent of usage remains on the most expensive frontier models even for trivial tasks, and that intelligent routing can yield order-of-magnitude savings by sending routine requests to cheaper models. Vendors report that although nominal per-token list prices have come down, the push to make systems more autonomous and the spread of higher-capability models have driven aggregate token consumption higher.
Practical changes are already appearing in teams. Developers who once had open access to high-tier models now face throttles, approval workflows and routing rules that send noncritical traffic to lower-cost models. Finance teams are adding billing alerts and breaking out token spend by project. Tool vendors are racing to provide dashboards, billing APIs and shared norms for measuring token efficiency. And some firms have quietly removed or scaled back features while they implement those controls.
The immediate management choice is blunt, and CEOs are framing it as "tokens or humans." Finance leads say their new job is deciding where to spend on models and where to preserve headcount. That trade-off plays out differently across sectors: some firms find net savings when AI replaces subscription services, others face runaway bills tied to high-volume experimentation and agentic workflows.
Behind the numbers is a governance puzzle. Token accounting has no widely accepted standard yet, so the Tokenomics Foundation and FinOps conversations aim to create common metrics and practices. Until those norms emerge, teams will juggle ad hoc quotas, routing logic and manual audits while vendors and open source projects add technical controls.
For now the effect on everyday users is simple: less open access. Teams that had been running unlimited or high-tier model usage have found new caps, billing alerts and routing controls applied to their accounts. The result is a period of retrenchment and measurement as companies seek to turn exploratory gains into disciplined, repeatable returns.
Related Articles
- OpenAI eyes token price cuts as customers ration AI use
- 90 investors back OpenAI and Anthropic, muting the rivalry
- Google to pay SpaceX $920M per month for AI compute
The Linux Foundation unveiled the Tokenomics Foundation in early June. Expect standards and vendor tools to follow as firms hunt for common token accounting and tighter controls.
This article was created with AI assistance.