Recent data from Ramp shows AI token spending for businesses jumped 572% from June 2025 to June 2026. The rapid rise means finance leaders can no longer treat AI costs as a one‑time line item; they behave more like infrastructure expenses and require ongoing oversight.
Why AI Token Costs Differ From Traditional SaaS
Most software budgets are predictable: SaaS licenses are seat‑based, cloud storage is capacity‑based, and both follow steady patterns. AI token costs, however, fluctuate with model choice, prompt length, call volume, and the teams using the technology. The median business sees month‑to‑month swings of about 58%, and 61% of companies experience variations of 40% or more.
Two‑Layer Approach: Engineering and Finance
Effective cost control starts with actions in the engineering stack—optimizing how models are used—and extends to the finance stack—providing visibility and accountability. Without a unified view, companies receive multiple invoices lacking context, making it difficult to pinpoint the source of cost spikes.
Common Drivers of Unexpected Spend
- Model drift upward: Teams often begin with a lower‑cost model for internal tools and later upgrade to premium versions for production without adjusting budgets.
- Long‑context inflation: Feeding large documents or full conversation histories into each request consumes tokens faster than anticipated.
- Multi‑model sprawl: Without clear policies, teams add new models incrementally, creating hidden contracts and cumulative spend.
These patterns compound because no single team owns full visibility; invoices arrive only after the spend has occurred.
Tracking and Allocating Costs
Finance systems must capture token consumption by business unit, project, cost center, and model. Distinguishing between cost of goods sold (COGS) and operating expenses (OpEx) is crucial. Tokens that power customer‑facing features belong in CGS, while internal‑tool usage should be recorded as OpEx. Mixing the two skews gross margin calculations and hampers unit‑economics modeling.
Practical Cost‑Reduction Tactics
- Prompt caching: Storing repeated system prompts or document contexts can cut input token costs by up to 90% (Anthropic).
- Model selection and routing: Direct simple tasks—such as classification or summarization—to smaller, cheaper models, reserving larger models for complex reasoning.
- Batch processing: Use asynchronous batch APIs for workloads that do not require instant responses; Anthropic’s Message Batches API can reduce costs by 50%.
- Spend limits by team: Implement soft monthly alerts, progressing to hard caps as tracking matures.
- Chargeback to cost centers: Allocate spend back to the teams that generated it, giving them a direct incentive to optimize.
- Anomaly alerting: Automated flags for spikes—e.g., a 40% month‑over‑month increase—provide early warnings for finance teams.
Building a Sustainable Framework
Start by pulling the last three months of AI spend by provider and mapping it to the responsible teams. Set realistic, soft budget targets based on current usage, then require a cost model for any new workload before it scales. Review spend monthly rather than quarterly to catch overruns early.
By combining engineering‑level efficiencies with finance‑level accountability, businesses can tame the volatility of AI token expenses and ensure that AI remains a strategic asset rather than an unchecked cost driver.
Original reporting: KRDO (Colorado Springs metro) — read the source article.