The most useful forecast is not the price of whatever model happens to be called frontier in 2029. It is the price of buying a fixed level of capability—the quality available from a strong model today.
That distinction changes the answer. The newest models will probably consume much of the industry's efficiency gain through longer reasoning, larger contexts, more tool calls, and more agent steps. Their sticker prices may look familiar. Meanwhile, today's intelligence should become much cheaper and faster.
My base case is that by July 2029, a model matching the capability of a strong July 2026 general-purpose model will cost $0.03–$0.50 per million output tokens and generate roughly 700–1,500 visible tokens per second in a single stream. The center of my range is below $0.10 and near 1,000 TPS.
That is a forecast, not an industry consensus. The ranges are deliberately wide because the historical improvements are extraordinary, the measurement is messy, and retail API prices reflect supply, demand, and strategy—not just the technical cost of inference.
Start with a July 2026 baseline
The cleanest current anchor is a pair of strong, general-purpose APIs with public list prices. OpenAI prices GPT-5.6 Luna at $1 per million input tokens and $6 per million output tokens. Artificial Analysis measures its public API at about 190 output tokens per second. Anthropic says Claude Sonnet 4.6 starts at $3 input and $15 output per million tokens.
So a reasonable July 2026 reference range for a strong model is:
| Metric | July 2026 reference |
|---|---|
| Input price | $1–$3 / 1M tokens |
| Output price | $6–$15 / 1M tokens |
| Single-stream output speed | About 180–200 TPS |
This is not a claim that Luna and Sonnet are identical. It is an observable price-and-speed bracket for models capable enough to anchor the forecast.
My forecast for today's capability
| Around | Input price / 1M | Output price / 1M | Typical single-stream TPS |
|---|---|---|---|
| July 2027 | $0.10–$0.70 | $0.60–$4 | 250–500 |
| July 2028 | $0.02–$0.20 | $0.12–$1.50 | 400–900 |
| July 2029 | $0.005–$0.07 | $0.03–$0.50 | 700–1,500 |
The central mental model behind those ranges is simple:
- Constant-capability price improves about 5× per year.
- Single-request generation speed improves about 2× per year.
- Global aggregate token capacity improves about 3× per year.
The first number is intentionally conservative relative to recent history. Epoch AI compared the cheapest models that cleared fixed benchmark thresholds and found price declines from 9× to 900× per year, depending on the task and threshold. GPT-4-level performance on GPQA fell about 40× annually. Epoch also warns that the fastest declines were concentrated in the most recent period and may not persist. Its study excluded reasoning models from the per-token analysis and blended input and output prices, so it should not be read as a literal promise about any one API bill.
Fivefold annual improvement is my haircut for slower future gains, benchmark saturation, provider margin, capacity constraints, and the possibility that the easy compression and serving wins arrive early.
Why frontier prices may barely fall
Constant capability and frontier capability are two different products.
OpenAI's July 2026 family makes the distinction visible inside one model generation: Luna costs $6 per million output tokens, Terra $15, and Sol $30. The provider is not only selling tokens. It is selling access to a higher capability tier. Google likewise says agent usage is billed for input, output, and intermediate reasoning tokens generated during agent loops.
When inference becomes more efficient, a lab has at least four choices: lower the price, increase margin, serve more demand, or spend more compute on each answer. Frontier products will use all four. More test-time reasoning and longer tool-using trajectories can absorb enormous efficiency gains without changing the visible length of the final response.
That is why I expect ordinary flagship output prices in 2029 to remain in roughly the same broad zone as today—single digits to tens of dollars per million tokens—with maximum-reasoning or priority tiers above it. The intelligence purchased at that price should be much greater.
Speed will rise, but latency will move elsewhere
The fastest useful models are already well beyond the strong-model baseline. Artificial Analysis currently measures Mercury 2 around 900–1,100 TPS as the live result moves over time, while Gemini 3.5 Flash-Lite is around 439 TPS. OpenAI has separately advertised limited GPT-5.6 Sol access on Cerebras at up to 750 TPS.
For the fastest useful inference-optimized models, my rough forecast is:
| Around | Fast-model single-stream TPS |
|---|---|
| 2027 | 1,500–3,000 |
| 2028 | 3,000–6,000 |
| 2029 | 5,000–10,000 |
These figures matter more for autonomous agents than for human chat. At 500 TPS, a 500-token answer streams in about one second. Beyond that, perceived speed is dominated by input processing, time to first token, reasoning time, tools, and the network. MLCommons therefore reports both throughput and latency constraints; its LLM benchmark work treats time to first token and time per output token as separate requirements.
Fast serving also costs more. In one SemiAnalysis InferenceX configuration, a DeepSeek-style workload cost about $0.56 per million output tokens at 50 TPS per user and about $4 at 125 TPS. A 2.5× speed increase cost roughly 7× more because the server used more parallel hardware per request and gave up batching efficiency. Future APIs will continue to expose a price-speed frontier, not one universal TPS number.
Supply is compounding, but demand may compound faster
The 3× capacity assumption is not purely a guess. Epoch estimates that global AI computing capacity grew 3.3× per year from 2022, with a 90% confidence interval of 2.7× to 4.1×. Its more detailed inference model estimates fixed-workload token capacity growing about 3.4× annually.
But the same analysis places potential token demand growth near 10× per year, with very high uncertainty. If demand outpaces deployed capacity, retail prices can fall more slowly than technical serving costs. Long-context agent workloads are especially exposed because prefill, memory bandwidth, and repeated context all consume capacity before visible generation begins.
This is the largest downside risk to the forecast. Better chips and software can make a token cheaper to produce while a shortage keeps the market price high.
Forecast cost per successful task, not just cost per token
An agent company should not put “token prices fall 5× per year” into its model and assume inference spending falls with it.
The useful operating equation is:
Inference spend = price at equivalent capability × tokens per successful task × number of tasks
The first term can fall 5× while the other two rise faster. Cheaper inference makes longer reasoning, multiple candidates, verification passes, computer use, and whole new classes of tasks economical. Agent software repeatedly discovers more expensive things to do with intelligence.
A practical planning case is:
| Driver | Planning assumption |
|---|---|
| Price per equivalent token | Down about 5× annually |
| Tokens consumed per successful task | Up 3–10× annually |
| Number of attempted agent tasks | Potentially up faster still |
| Total inference spend | Flat or increasing |
The headline is therefore not “tokens become irrelevant.” It is: today's intelligence becomes nearly free while frontier intelligence and total inference demand remain expensive.
That outcome would accelerate the service-buying loops described in Agentic Commerce Is Becoming a Services Market: agents will be able to search longer, compare more suppliers, negotiate more iterations, and verify more outputs for the same unit cost. It also raises the value of outcome evidence. More cheap agent actions mean more transactions that need to be judged after the fact.
What to track from here
No single source answers this forecast. These are the indicators I would follow:
- Artificial Analysis: live first-party and provider comparisons for price, output TPS, time to first token, token use, and cost per benchmark task.
- Epoch AI: price at fixed performance, aggregate compute capacity, inference economics, and supply-demand forecasts.
- SemiAnalysis InferenceX: tokens per second per GPU, tokens per second per user, batching, quantization, speculative decoding, and modeled serving cost.
- MLCommons MLPerf Inference: standardized hardware comparisons under explicit accuracy and latency constraints.
- Official pricing pages: input, output, cached input, batch, priority, long-context, and reasoning-token charges from the labs themselves.
I would record a monthly snapshot for three capability bands—frontier, strong, and fast-cheap—and keep price, cost per fixed task, TPS, time to first token, and tokens per task separate. That small discipline prevents a faster but more verbose reasoning model from looking cheaper simply because its output-token sticker price fell.
The forecast will be wrong in its exact numbers. The durable part is the measurement frame: hold capability constant, price the whole successful task, and keep frontier progress separate from efficiency progress.
Primary sources
- Epoch AI: LLM inference prices have fallen rapidly but unequally across tasks
- Epoch AI: Is a compute crunch coming?
- Epoch AI: Global AI computing capacity is doubling every 7 months
- SemiAnalysis InferenceX v2
- MLCommons: MLPerf Inference v5.0 language-model methodology
- Artificial Analysis: GPT-5.6 Luna
- OpenAI: GPT-5.6 Luna pricing
- Anthropic: Claude Sonnet 4.6
- Google: Gemini API pricing
