# AI overspend statistics 2026

Last updated 2026-08-06. Every number sourced — cite with a link back to https://cutmyaispend.com/stats.

- **79%** of enterprises overspent on AI in 2026. Sapio Research survey (Feb 2026, commissioned by DoiT) of 500 finance leaders at 1,000+ employee organizations across the US and UK. (Source: DoiT / Sapio Research — https://www.doit.com/blog/ai-spending-survey)
- **73%** of enterprise agentic-AI implementations went over budget. Review of 127 enterprise agentic AI implementations; some exceeded original estimates by more than 2.4×, burning ~$2.3M in unanticipated costs. (Source: BERI AI FinOps analysis, 2026 — https://www.beri.net/article/ai-finops-2026-73-percent-blow-budget-cfo-fix)
- **89%** of organizations that call their FinOps "very mature" still had AI cost overruns. Mean overspend in this segment reached 30.9% — the highest of any group studied. Mature orgs run bigger AI programs and actually detect their overruns. (Source: DoiT / Sapio Research, 2026 — https://www.doit.com/blog/ai-spending-survey)
- **98%** of FinOps teams now manage AI spend — up from 31% two years ago. AI workloads have grown to ~18% of cloud budgets at AI-forward companies, up from 4% in 2023. (Source: FinOps X 2026 — https://www.usage.ai/blogs/finops/ai-ml-cost/finops-x-2026-takeaways)
- **Up to 90%** off cached input tokens with provider prompt caching. Anthropic prompt caching discounts cached tokens up to 90%; OpenAI applies ~50% automatically. Production cache-hit rates cluster at 50–80%. (Source: Provider pricing docs; production reports — https://www.digitalapplied.com/blog/prompt-caching-2026-cut-llm-costs-engineering-guide)
- **Up to 98%** cost reduction demonstrated by LLM cascade routing at matched quality. FrugalGPT-style cascades (Stanford) send queries to cheap models first and escalate only when needed. (Source: Stanford FrugalGPT research — https://neuraltrust.ai/blog/llm-cost-reduction-guide)
- **3–8×** output tokens cost more than input tokens (median ~4:1). Long responses are billed at the premium rate — output-length control is one of the cheapest savings available. (Source: Cross-provider pricing analysis, 2026 — https://neuraltrust.ai/blog/llm-cost-reduction-guide)
- **30–70%** of redundant API calls eliminated by semantic caching. Matching paraphrased queries against previously answered ones skips the model call entirely in FAQ-heavy workloads. (Source: NeuralTrust LLM cost reduction guide, 2026 — https://neuraltrust.ai/blog/llm-cost-reduction-guide)
- **50%** flat discount from batch APIs at every major provider. OpenAI, Anthropic, and Google all price async batch endpoints at ~half the synchronous rate. (Source: Provider pricing docs, 2026 — https://www.getmaxim.ai/articles/reduce-llm-cost-and-latency-a-comprehensive-guide-for-2026/)
- **~10×** cheaper per agent task with a persistent context/memory layer. Mitosis Labs reports ~1/10th cost and 98% fewer hallucinations when agents query an indexed memory graph instead of re-ingesting raw data each run. (Source: Mitosis Labs — https://mitosislabs.ai)

---
Canonical: https://cutmyaispend.com/stats