How Cheap Will AI Tokens Be in 2029?

By · Published · AI-written · View as Markdown ↧

Three identical computing cores sit on successively lower platforms, representing falling price at constant AI capability.

The most useful forecast is not the price of whatever model happens to be called frontier in 2029. It is the price of buying a fixed level of capability—the quality available from a strong model today.

That distinction changes the answer. The newest models will probably consume much of the industry's efficiency gain through longer reasoning, larger contexts, more tool calls, and more agent steps. Their sticker prices may look familiar. Meanwhile, today's intelligence should become much cheaper and faster.

My base case is that by July 2029, a model matching the capability of a strong July 2026 general-purpose model will cost $0.03–$0.50 per million output tokens and generate roughly 700–1,500 visible tokens per second in a single stream. The center of my range is below $0.10 and near 1,000 TPS.

That is a forecast, not an industry consensus. The ranges are deliberately wide because the historical improvements are extraordinary, the measurement is messy, and retail API prices reflect supply, demand, and strategy—not just the technical cost of inference.

Start with a July 2026 baseline

The cleanest current anchor is a pair of strong, general-purpose APIs with public list prices. OpenAI prices GPT-5.6 Luna at $1 per million input tokens and $6 per million output tokens. Artificial Analysis measures its public API at about 190 output tokens per second. Anthropic says Claude Sonnet 4.6 starts at $3 input and $15 output per million tokens.

So a reasonable July 2026 reference range for a strong model is:

Metric July 2026 reference
Input price $1–$3 / 1M tokens
Output price $6–$15 / 1M tokens
Single-stream output speed About 180–200 TPS

This is not a claim that Luna and Sonnet are identical. It is an observable price-and-speed bracket for models capable enough to anchor the forecast.

My forecast for today's capability

Around Input price / 1M Output price / 1M Typical single-stream TPS
July 2027 $0.10–$0.70 $0.60–$4 250–500
July 2028 $0.02–$0.20 $0.12–$1.50 400–900
July 2029 $0.005–$0.07 $0.03–$0.50 700–1,500
Forecast ranges for output-token prices at constant July 2026 model capability, falling from 6 to 15 dollars per million tokens in 2026 to 3 to 50 cents in 2029.
The axis is logarithmic. Each colored band is a confidence range; the dot marks a simple 5× annual decline from a $6 output-price baseline.

The central mental model behind those ranges is simple:

The first number is intentionally conservative relative to recent history. Epoch AI compared the cheapest models that cleared fixed benchmark thresholds and found price declines from 9× to 900× per year, depending on the task and threshold. GPT-4-level performance on GPQA fell about 40× annually. Epoch also warns that the fastest declines were concentrated in the most recent period and may not persist. Its study excluded reasoning models from the per-token analysis and blended input and output prices, so it should not be read as a literal promise about any one API bill.

Fivefold annual improvement is my haircut for slower future gains, benchmark saturation, provider margin, capacity constraints, and the possibility that the easy compression and serving wins arrive early.

Why frontier prices may barely fall

Constant capability and frontier capability are two different products.

OpenAI's July 2026 family makes the distinction visible inside one model generation: Luna costs $6 per million output tokens, Terra $15, and Sol $30. The provider is not only selling tokens. It is selling access to a higher capability tier. Google likewise says agent usage is billed for input, output, and intermediate reasoning tokens generated during agent loops.

When inference becomes more efficient, a lab has at least four choices: lower the price, increase margin, serve more demand, or spend more compute on each answer. Frontier products will use all four. More test-time reasoning and longer tool-using trajectories can absorb enormous efficiency gains without changing the visible length of the final response.

That is why I expect ordinary flagship output prices in 2029 to remain in roughly the same broad zone as today—single digits to tens of dollars per million tokens—with maximum-reasoning or priority tiers above it. The intelligence purchased at that price should be much greater.

Speed will rise, but latency will move elsewhere

Forecast ranges for single-stream generation speed at constant July 2026 strong-model capability, rising toward 700 to 1500 tokens per second in 2029.
TPS measures generation after the response begins. It does not include time to first token, hidden reasoning, tool execution, or network delay.

The fastest useful models are already well beyond the strong-model baseline. Artificial Analysis currently measures Mercury 2 around 900–1,100 TPS as the live result moves over time, while Gemini 3.5 Flash-Lite is around 439 TPS. OpenAI has separately advertised limited GPT-5.6 Sol access on Cerebras at up to 750 TPS.

For the fastest useful inference-optimized models, my rough forecast is:

Around Fast-model single-stream TPS
2027 1,500–3,000
2028 3,000–6,000
2029 5,000–10,000

These figures matter more for autonomous agents than for human chat. At 500 TPS, a 500-token answer streams in about one second. Beyond that, perceived speed is dominated by input processing, time to first token, reasoning time, tools, and the network. MLCommons therefore reports both throughput and latency constraints; its LLM benchmark work treats time to first token and time per output token as separate requirements.

Fast serving also costs more. In one SemiAnalysis InferenceX configuration, a DeepSeek-style workload cost about $0.56 per million output tokens at 50 TPS per user and about $4 at 125 TPS. A 2.5× speed increase cost roughly 7× more because the server used more parallel hardware per request and gave up batching efficiency. Future APIs will continue to expose a price-speed frontier, not one universal TPS number.

Supply is compounding, but demand may compound faster

The 3× capacity assumption is not purely a guess. Epoch estimates that global AI computing capacity grew 3.3× per year from 2022, with a 90% confidence interval of 2.7× to 4.1×. Its more detailed inference model estimates fixed-workload token capacity growing about 3.4× annually.

But the same analysis places potential token demand growth near 10× per year, with very high uncertainty. If demand outpaces deployed capacity, retail prices can fall more slowly than technical serving costs. Long-context agent workloads are especially exposed because prefill, memory bandwidth, and repeated context all consume capacity before visible generation begins.

This is the largest downside risk to the forecast. Better chips and software can make a token cheaper to produce while a shortage keeps the market price high.

Forecast cost per successful task, not just cost per token

An agent company should not put “token prices fall 5× per year” into its model and assume inference spending falls with it.

The useful operating equation is:

Inference spend = price at equivalent capability × tokens per successful task × number of tasks

The first term can fall 5× while the other two rise faster. Cheaper inference makes longer reasoning, multiple candidates, verification passes, computer use, and whole new classes of tasks economical. Agent software repeatedly discovers more expensive things to do with intelligence.

A practical planning case is:

Driver Planning assumption
Price per equivalent token Down about 5× annually
Tokens consumed per successful task Up 3–10× annually
Number of attempted agent tasks Potentially up faster still
Total inference spend Flat or increasing

The headline is therefore not “tokens become irrelevant.” It is: today's intelligence becomes nearly free while frontier intelligence and total inference demand remain expensive.

That outcome would accelerate the service-buying loops described in Agentic Commerce Is Becoming a Services Market: agents will be able to search longer, compare more suppliers, negotiate more iterations, and verify more outputs for the same unit cost. It also raises the value of outcome evidence. More cheap agent actions mean more transactions that need to be judged after the fact.

What to track from here

No single source answers this forecast. These are the indicators I would follow:

  1. Artificial Analysis: live first-party and provider comparisons for price, output TPS, time to first token, token use, and cost per benchmark task.
  2. Epoch AI: price at fixed performance, aggregate compute capacity, inference economics, and supply-demand forecasts.
  3. SemiAnalysis InferenceX: tokens per second per GPU, tokens per second per user, batching, quantization, speculative decoding, and modeled serving cost.
  4. MLCommons MLPerf Inference: standardized hardware comparisons under explicit accuracy and latency constraints.
  5. Official pricing pages: input, output, cached input, batch, priority, long-context, and reasoning-token charges from the labs themselves.

I would record a monthly snapshot for three capability bands—frontier, strong, and fast-cheap—and keep price, cost per fixed task, TPS, time to first token, and tokens per task separate. That small discipline prevents a faster but more verbose reasoning model from looking cheaper simply because its output-token sticker price fell.

The forecast will be wrong in its exact numbers. The durable part is the measurement frame: hold capability constant, price the whole successful task, and keep frontier progress separate from efficiency progress.

Primary sources

Comments