Blog

Why Your AI Inference Bill Keeps Growing, Even as the Price Per Token Falls

By NorthStar Group · August 24, 2026

The price of an AI call is dropping fast. Your AI bill is still going up. Both are true at once, and the gap between them is where budgets quietly break. AI inference costs are projected to fall more than 90% by 2030, yet at least half of GenAI projects are expected to overrun their budgets through 2028. Cheaper tokens, far more of them, a bigger bill. If your pilot looked cheap and your operating cost is climbing without a clear reason, here's why.

Key takeaways

Inference is 70 to 90% of AI compute spend, and it never stops.

Half of GenAI projects will overrun their budgets through 2028.

Prices fall 90% by 2030; consumption grows faster.

The bill surprises you because nobody owns the run cost.

The mistake: budgeting for the pilot, not the production bill

Inference is what your AI system costs every time it answers: a recurring, per-call cost. A pilot serves a handful of users a few hundred times a day, and that bill is small enough to ignore. Which is exactly the problem. The pilot number becomes the mental anchor, then production arrives and the anchor was never real.

The mechanics are usually some mix of these:

01

A negligible per-call cost, multiplied by real volume. A cost that rounds to zero for one user doesn't round to zero across ten thousand.

02

Agentic workloads that call the model many times per task. Cost scales with steps, not users.

03

Retries, retrieval, and long context that nobody metered.

04

No visibility into which feature or team is spending. The invoice arrives as one number, so there's nothing to optimize. That vacancy is the subject of who owns your AI run cost.

The pilot proved it works. It didn't prove it's affordable at scale.

Why does the bill grow even as prices fall?

Because you're buying far more tokens than the price is dropping. Cheaper inference reads like a saving, but in practice it's permission to use more, and organizations take that permission faster than the price falls. Economists call this the Jevons paradox: make a resource cheaper and total consumption rises enough to more than absorb the saving. Waiting for the price to fall on its own is a plan to spend more.

What does an unmanaged inference bill actually cost?

Most GenAI projects overrun, and the cause is architectural. Through 2028, at least 50% of GenAI projects will overrun their budgeted costs, driven by poor architectural choices and no visibility into how cost scales. That's a planning failure, which means it's preventable.

Token consumption is on track to rival payroll. AI coding token costs are forecast to surpass the average developer's salary by 2028, with individual developers already reported at $20,000 of tokens in a month. It has happened at named-company scale: Uber exhausted its entire 2026 AI budget by April, after agentic coding tools reached roughly 5,000 engineers and adoption jumped from 32% to 84% in one month. Meta imposed an internal token cap the same quarter.

Agentic AI makes the curve steeper, and cancellations follow. Agentic systems multiply calls per task and therefore cost per task. Over 40% of agentic AI projects are forecast to be canceled by the end of 2027, with escalating costs among the leading causes. Many were delivering value, killed anyway because the cost to run them was never owned. Pricing a full task at volume is one of the five screens in spotting agent-washing before you sign.

The contrast: unmanaged vs. governed inference

✕ The unmanaged way

Budget anchored on the pilot's cost

✓ The governed way

Cost modeled at production volume before launch

✕ The unmanaged way

One invoice, no idea what drives it

✓ The governed way

Spend attributed by feature, team, and workflow

✕ The unmanaged way

Biggest model for every call

✓ The governed way

Right-sized model per task, cheapest that clears the bar

✕ The unmanaged way

Agent free to call itself without limit

✓ The governed way

Step and cost budgets enforced per task

✕ The unmanaged way

Someone owns the build

✓ The governed way

Someone owns the run cost, with a number attached

The right column means knowing what each call costs and buying only the ones worth buying, using AI just as much.

What we see in the field

Across our AI-enablement and data pods, the inference bill is almost never high because the workload is high. It's high because the architecture was never designed for cost. The first move is to instrument and attribute the spend, so the invoice stops being one opaque number and starts pointing at a cause.

The same few culprits turn up: most token spend going to a handful of workflows, oversized models doing work a smaller model would do for a fraction of the cost, retrieval pulling more context than the answer needs, and agent loops that retry without a limit. Fixing the architecture typically takes token spend down 40 to 60% at the same output. The number drops because the waste is gone, not because the AI does less.

What this means for you

Your time

Modeling and instrumenting inference cost up front costs you days. The alternative costs you quarters: explaining a growing bill, then re-architecting under pressure once finance flags it.

Your money

The unit price is falling and your total is still rising, so your lever is architecture, not waiting for cheaper tokens. Governed inference routinely cuts spend 40 to 60% for the same output.

Your risk

An AI feature canceled for cost after it shipped and worked is a reputational hit with your name on the sponsor line. Owning the run cost is how your working system stays funded instead of being killed by its own invoice.

Every call is cheap. Your total isn't, and right now nobody's on the hook for it.

Frequently asked questions

Why are my AI inference costs growing unexpectedly?

Because production looks nothing like the pilot. A per-call cost that rounds to zero for a few users multiplies across real volume, agent calls, oversized prompts, and retries. The bill grows fastest where nobody can see what's driving it.

Why does my AI bill go up when the price per token is falling?

Because consumption grows faster than the price drops. Inference is projected to cost over 90% less by 2030, but cheaper tokens invite far more usage, so the total keeps climbing even as each call gets cheaper.

How do you reduce AI inference costs without cutting features?

Instrument spend so you can see what drives it, right-size the model per task, cut wasted context, cache repeated calls, and put cost budgets on agent loops. The savings come from removing waste, which is why governed inference commonly cuts spend 40 to 60% at the same output.

See where your AI run cost is exposed

If your AI bill is climbing and the invoice doesn't tell you why, a 30-minute inference-cost review gets you a whiteboard read on where the spend is concentrated and what governance would change. No obligation.

AI made inference fast and cheap per call. NorthStar makes it accountable.

A vetted, dedicated pod that owns the outcome inside your systems, including what it costs to run, without the Big-4 bill or the offshore babysitting.

Sources

Gartner, "Gartner Predicts AI Coding Costs Will Surpass Average Developer's Salary by 2028 as Token Consumption Surges" (June 24, 2026). gartner.com

Gartner, "10 Best Practices for Optimizing Generative and Agentic AI Costs": through 2028, at least 50% of GenAI projects will overrun their budgeted costs.

Gartner, trillion-parameter inference cost to fall over 90% by 2030 (March 25, 2026). gartner.com

Gartner, over 40% of agentic AI projects canceled by end of 2027 (June 25, 2025). gartner.com

Fortune, "Uber burned through its entire 2026 AI budget in four months" (May 26, 2026): ~5,000 engineers, adoption 32% to 84% in one month. fortune.com

Inference share of AI compute spend (70 to 90% over a model's production life): industry-consensus estimate from AI infrastructure and FinOps analyses. Directional; verify before quoting in a client-facing asset.

Keep Reading

Scroll to Top