Signals Inbox·July 28, 2026·Frontier AI
How much will a token cost in 2027?
A capable mainstream AI model should cost about $1 per million blended tokens in 2027, but the real market will stretch from almost-free cached input to premium reasoning that still costs tens of dollars per million tokens.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →A mainstream AI token should cost roughly one-millionth of a dollar in 2027. Our central estimate is $0.50 per million input tokens and $3 per million output tokens, equal to about $1.13 per million tokens on a three-to-one input-output mix.
The cheapest listed rate will matter less than it appears. Caching, batch processing and model routing can remove most input cost, while retries, reasoning tokens and tool calls can make a cheap model more expensive per completed task.
Today's frontier capability should become three to ten times cheaper. The newest frontier model probably will not: providers will use cheaper compute to open another premium tier with longer reasoning, larger contexts and better agent performance.
Token prices will fall faster than many AI bills. Agents can make hundreds of calls, resend large contexts and consume far more tokens than chat, so companies may spend more overall even while every million tokens gets cheaper.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ1What does “the price of an AI token” actually mean?
There will be no single AI token price in 2027 because providers already sell several very different kinds of tokens.
A model charges one rate for text it reads and another for text it writes. Repeated prompt content may receive a large cache discount. Slow batch jobs can cost half as much as real-time requests. Some providers also charge more for priority processing, long prompts, data residency, web searches or tool calls.
Tokenizers add another complication. The same paragraph can produce different token counts across models and languages. DeepSeek estimates roughly 0.3 tokens per English character and 0.6 tokens per Chinese character with its tokenizer, while other models split the text differently. A million tokens from two providers may represent different amounts of language.
For this article, we use three prices. Economy pricing covers cheap models used for routing, extraction and simple generation. Mainstream pricing covers capable production models used for support, coding and document work. Frontier pricing covers the strongest models available for difficult reasoning and agent tasks. We also separate input from output because combining them too early creates a fake sense of precision.
Q2Why is the 2027 AI token price so hard to predict?
The 2027 AI token price is hard to predict because cheaper computing and more demanding models are pulling prices in opposite directions.
Hold capability steady and the price collapse looks spectacular. Follow the strongest model available each year and the decline looks much smaller. Providers use cheaper computing to add longer reasoning, larger context windows, faster service and better tool use, then open a new premium tier.
OpenAI’s current range shows the problem clearly. GPT-5.6 Luna costs $1 per million input tokens and $6 per million output tokens. Terra costs $2.50 and $15. Sol costs $5 and $30. All three belong to the same model family, yet the most expensive tier costs five times more than the cheapest.
Customers react differently too. One company may move an old workload to a cheaper model and save money. Another may spend the saving on agents that make hundreds of calls. Posted rates are easier to forecast than the final bill.
Q3How much does an AI token cost right now?
AI token prices currently span more than a hundredfold, even before enterprise discounts enter the picture.
At the low end, full production APIs now start near $0.15 per million input tokens and roughly $0.30 to $0.60 for output. These models can reason, call tools and support agents.
The middle of the market currently sits around $1 to $1.50 for input and $6 to $7.50 for output. Premium models remain expensive: Claude Opus 5 costs $5 and $25, GPT-5.6 Sol costs $5 and $30, and Claude Fable 5 reaches $10 and $50.
Anthropic currently prices Claude Sonnet 5 at an introductory $2 for input and $10 for output, with the standard rate rising to $3 and $15 later in 2026. That scheduled increase shows providers still expect a strong mid-to-high tier to hold its price as cheaper models improve.
We calculated the blended column below with three input tokens for every output token. Real applications vary widely, but using the same ratio makes the gap easier to see.
Current AI API token prices
| Model | Input per 1M | Output per 1M | 3:1 blended price |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | $0.18 |
| Mistral Small 4 | $0.15 | $0.60 | $0.26 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.85 |
| GPT-5.6 Luna | $1.00 | $6.00 | $2.25 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $3.00 |
| Claude Sonnet 5, introductory rate | $2.00 | $10.00 | $4.00 |
| Claude Opus 5 | $5.00 | $25.00 | $10.00 |
| GPT-5.6 Sol | $5.00 | $30.00 | $11.25 |
| Claude Fable 5 | $10.00 | $50.00 | $20.00 |
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q4Have AI token prices really collapsed?
Yes. The price of yesterday’s intelligence has collapsed, even though the newest frontier models still carry premium rates.
Stanford’s 2025 AI Index tracked the cost of reaching roughly GPT-3.5 performance on the MMLU benchmark. It fell from about $20 per million tokens in late 2022 to $0.07 by late 2024, a reduction of more than 280 times in around 18 months.
Epoch AI reached a similar conclusion across six benchmarks. Depending on the task and target score, it measured annual capability-adjusted price declines ranging from ninefold to 900-fold. The range is huge because a benchmark can suddenly become cheap once smaller models master it.
List prices tell a less dramatic story at the top. GPT-4 launched at $30 per million input tokens and $60 per million output tokens. GPT-5 later fell to $1.25 and $10. The current GPT-5.6 flagship has moved back to $5 and $30 because buyers are getting a newer class of model with longer context, stronger reasoning and broader tool use.
Anthropic provides another clean example. Claude Opus 4 cost $15 for input and $75 for output. Opus 5 is currently $5 and $25, a two-thirds reduction within the premium tier. Anthropic then introduced Fable 5 at $10 and $50, creating another tier above it.
That staircase should continue: older capabilities slide down the price ladder, then each new frontier reopens the premium tier.
Q5Will today’s best AI be cheap by 2027?
Yes. Today’s frontier-level AI should cost three to ten times less to use by the end of 2027.
We are forecasting the cost of doing today’s work, regardless of which model name leads the menu in 2027. A task that currently needs a $5-input, $25-output model may run well on a 2027 mid-tier model costing a fraction of that amount.
The change is already visible inside current model families. Google sells Gemini 3.5 Flash-Lite at $0.30/$2.50 and its stronger Flash tier at $1.50/$7.50. Anthropic’s Haiku 4.5 costs $1/$5, far below Opus 5 at $5/$25.
Each new low-cost tier takes on work that recently needed a flagship. Once cheaper rivals can handle the same customer task, the old premium gets hard to defend.
Routine classification, extraction, translation and coding could see tenfold declines. Hard research, long-running agents and high-stakes reasoning will probably land nearer the threefold end of the range.
Q6Why does the model charge more to write than to read?
Output tokens will still cost several times more than input tokens in 2027 because writing a response keeps the model working one step at a time.
A model can read many prompt tokens at once. Writing is slower: it produces one token, checks the growing sentence, then produces the next. The longer the answer, the longer the chip stays tied up.
Current prices show a clear pattern. Across the main Claude and Gemini models, output costs about five times as much as input. OpenAI’s mainstream tiers are closer to six times, while several low-cost models sit near four times.
Reasoning makes the output bill larger. Google counts thinking tokens as output, and other providers also charge for hidden or visible reasoning through their token accounting. A short answer can carry a much larger amount of paid computation behind it.
Faster decoding and better chips will lower both rates. Mainstream output should still cost roughly five or six times more than standard input in 2027.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxOpenAI's product lead says future AI won’t fit on your laptop
Mistral and Saudi Arabia are building frontier AI in Arabic
Open-weight models jumped from 28% to 62% on Vercel
OpenCode is giving away Ox Alpha, with 1M-token context
Stripe tells investors the Singularity began on January 1, 2026
Anthropic adds $18B in annualized revenue, reaching $65B
OpenAI’s longtime COO Brad Lightcap leaves to build something new
Anthropic hires a policy veteran to navigate its Trump standoff
Alibaba unveils Qwen3.8-Max, China’s latest shot at frontier AI
OpenAI says Astra found 10 breakthroughs humans missed for decades
Thinking Machines opens full weights to a 276B sparse model
OpenAI slashes GPT-5.6 Luna pricing by 80% overnight
Q7How much can caching and batch jobs cut the bill?
Caching and batch processing can already remove 50% to more than 95% of token costs from the right workload.
Batch pricing is the simplest saving. Google, Anthropic and Mistral currently offer 50% discounts for eligible asynchronous jobs. A company processing yesterday’s invoices, product descriptions or evaluation datasets can accept a delayed response and pay half the standard rate.
Caching goes further when the same long prompt appears repeatedly. OpenAI charges one-tenth of the normal input rate for cached GPT-5.6 tokens. Anthropic also charges 10% of its base input rate for cache hits. Google’s cached Gemini input commonly costs one-tenth of the standard rate, although storage can add a separate charge.
DeepSeek is especially aggressive. V4 Flash falls from $0.14 per million uncached input tokens to $0.0028 for a cache hit. Reprocessing one billion tokens would cost $140 without the cache and $2.80 with it.
Real savings depend on prompt design. The repeated material needs to remain stable, and requests must hit the provider’s cache. Teams that keep instructions, tool definitions and reference documents at the start of the prompt have a much better chance of getting the discount.
By 2027, large buyers may pay almost nothing for repeated input. Generated output, tool calls and failed agent steps will be the expensive part.
Common ways to reduce AI token costs
| Pricing method | Typical reduction now | Best fit |
|---|---|---|
| Batch processing | About 50% | Evaluations, document jobs and offline generation |
| Standard prompt cache | About 90% on repeated input | Stable instructions, repeated documents and long system prompts |
| Model routing | Often 50% to 90% | Easy requests sent to cheap models, with hard requests escalated |
Q8Can the cheaper model still cost more in the end?
Yes. A model with a cheaper listed token rate can produce a larger final bill when it thinks longer, retries more often or fails the task.
A 2026 study led by researchers from Stanford, Berkeley and Microsoft compared eight reasoning models across nine tasks. In 21.8% of model-pair comparisons, the model with the lower advertised price ended up costing more. The largest reversal reached 28 times.
Thinking-token use caused most of the mismatch. One model could consume 900% more reasoning tokens than another on the same query. Repeated runs of the same model on the same problem also varied by as much as 9.7 times.
Imagine one coding model charges $5 per million output tokens and burns through 200,000 tokens across planning and corrections. The cost is $1. A second model charges $15 but succeeds with 40,000 tokens, costing $0.60. The expensive rate wins because the model finishes efficiently.
Providers now let users dial reasoning effort up or down. Turning it down can save money and hurt accuracy. The only reliable comparison is a test on the company’s own tasks, including retries and failure rates.
Q9Will DeepSeek and open models drag every token price down?
DeepSeek and open-weight models will crush prices in the economy tier, while premium models keep charging for reliability and maximum capability.
As seen above, DeepSeek now sets the public floor for API pricing. Mistral Small 4 adds another kind of pressure: it costs $0.15/$0.60, uses an Apache 2.0 license and supports a 256,000-token context window. Developers can use it for routing, extraction, translation, coding assistance and many agent steps.
The pressure reaches larger providers even when customers stay put. OpenAI, Google and Anthropic need inexpensive tiers that remain competitive enough to prevent easy workloads from leaving. They can respond through smaller models, caching, batch discounts and more efficient token use.
Premium pricing survives when the higher-priced model saves human time or avoids expensive mistakes. Enterprises may also pay for stronger uptime, security controls, support, regional processing and stable model versions. A model that finishes a complex task in one run can beat a cheaper rival that needs several attempts.
The floor will keep falling. The premium tier still has enough buyers to survive, so the wide price spread is not going away in 2027.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q10Will NVIDIA Rubin make AI tokens ten times cheaper?
On selected workloads, NVIDIA Rubin and newer inference software can cut the underlying cost per token by close to ten times. Retail prices will fall much less.
NVIDIA says its Vera Rubin platform can deliver up to ten times lower inference-token cost than Blackwell for long-context, reasoning-heavy workloads. That figure comes from a favorable system configuration and applies only to selected workloads. Deployment will also take time because providers need new servers, networking, power and cooling.
Software adds another layer. Mixture-of-experts models wake up only part of the network for each token. Quantization uses lighter-weight numbers. Speculative decoding lets a small model draft text that a larger model quickly checks. Better batching keeps expensive GPUs busy.
Power shortages and grid delays will absorb some of the saving. Providers can charge extra for guaranteed speed during busy periods, while moving flexible jobs to cheaper capacity. Google’s current menu already separates standard, batch, flex and priority inference, with priority Gemini 3.6 Flash costing 80% more than standard.
Providers can also spend the saving on stronger models while leaving the rate card roughly where it is. Longer reasoning, larger prompts and more tool calls are all ways to consume cheaper compute. That is why our 2027 retail forecast falls less than the underlying infrastructure cost.
Q11When is self-hosting actually cheaper than an API?
Self-hosting becomes cheaper when traffic is large, steady and predictable enough to keep the GPUs busy.
A 2026 study of H100 inference economics found that effective output cost on identical hardware ranged from $0.21 to $15.25 per million tokens. Utilization created almost the entire gap. Near-idle deployments suffered penalties of up to 36 times because the company kept paying for hardware while very few requests arrived.
Large platforms can make self-hosting work. They have continuous traffic, infrastructure teams and enough volume to spread engineering costs across billions of tokens. Open-weight models also give them more control over privacy, latency and model behavior.
Smaller companies often underestimate the hidden costs. They need spare capacity for traffic spikes, monitoring, upgrades, security, failover and engineers who understand the serving stack. A managed API turns those fixed costs into a variable bill and gives immediate access to newer models.
The break-even point differs sharply by workload. Self-hosting should win more often in 2027 for stable, high-volume jobs. APIs will remain attractive for unpredictable traffic and frontier models that change every few months.
Q12Will AI agents still be expensive when tokens get cheaper?
Yes. AI agents can stay expensive because a single task may trigger hundreds of model calls and repeatedly resend huge prompts.
One-million-token context windows make that easy. OpenAI, Anthropic and DeepSeek now offer them on major models, allowing an application to read large codebases, document collections or long histories. Sending the same one-million-token prompt 100 times during an agent run creates 100 million input tokens before the final answer appears.
Agents multiply calls as well as context. They plan, search, open files, use tools, read the results, correct errors and try again. A 2026 study of coding agents found that agentic tasks consumed around 1,000 times more tokens than code chat or isolated code reasoning. Runs on the same task could differ by as much as 30 times, and the study found no reliable link between extra token use and accuracy.
Teams can control the waste by retrieving only relevant files, caching stable instructions, compressing old conversations and using cheap models for routine steps.
Even with those improvements, total token demand should grow faster than token prices fall. A company may pay less for each million tokens in 2027 and still spend more overall because it asks models to do far more work.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ13Should companies stop comparing models by token price?
For serious workloads, yes. Token price matters less than the cost of a successful task.
A support team cares about the cost per resolved ticket. A software team cares about the cost per accepted code change. A document processor cares about the cost per correctly completed file. Cheap tokens have little value when they create extra retries, human review or customer complaints.
The calculation should include input, output, hidden reasoning, tool calls and failure rates. Latency belongs in the comparison too. A batch job may halve token cost, while a customer-facing assistant needs an immediate response. Priority service can cost more and still create more value.
Model routing is already becoming standard. Easy tasks go to a low-cost model, difficult cases move to a stronger one, and the company measures the blended cost and success rate across the whole workflow.
Q14What token budget should companies use for 2027?
Companies should plan around $0.50 per million input tokens and $3 per million output tokens for a capable mainstream model in 2027.
With three input tokens for every output token, that equals about $1.13 per million blended tokens. A support exchange using 2,000 input tokens and 500 output tokens would cost roughly $0.0025. One dollar would cover around 400 such exchanges before search, storage, tool and application costs.
A larger agent task using 500,000 input tokens and 100,000 output tokens would cost about $0.55 at the same rates. A frontier model priced at $4 for input and $25 for output would charge $4.50 for that raw volume. The frontier option may still win when it avoids retries or completes a task the mainstream model cannot handle.
Usage cannot stay buried inside the price forecast. If agent adoption multiplies token volume fivefold, even a 60% rate cut leaves the bill twice as high. Budgeting needs separate cases for ordinary chat, growing agent use and heavy long-context automation.
The cheapest setup will mix models: inexpensive ones handle predictable work, a stronger model receives the hard cases, repeated context is cached and slow jobs run in batches.
Q15So how much will one AI token cost in 2027?
Our answer is about $1 per million blended tokens for mainstream AI in 2027, or roughly one-millionth of a dollar per token.
That estimate comes from $0.50 per million normal input tokens and $3 per million output tokens for a capable production model. Using a three-to-one input-output mix, the blended rate reaches $1.13 per million tokens, or $0.00000113 for one token.
Economy models should reach roughly $0.03 to $0.15 per million input tokens and $0.15 to $0.75 per million output tokens. Cached or batched input can fall below $0.01 in favorable cases.
Frontier models should stay around $2 to $6 for input and $12 to $35 for output. A small premium tier may exceed those ranges when it offers unusually deep reasoning, very fast service or specialist performance.
Expected AI token prices in 2027
| 2027 model category | Input per 1M | Output per 1M | 3:1 blended price |
|---|---|---|---|
| Cached or batch economy | $0.005–$0.08 | $0.10–$0.50 | $0.03–$0.19 |
| Standard economy | $0.03–$0.15 | $0.15–$0.75 | $0.06–$0.30 |
| Mainstream production | $0.20–$1.00 | $1.00–$6.00 | $0.40–$2.25 |
| Frontier general-purpose | $2.00–$6.00 | $12.00–$35.00 | $4.50–$13.25 |
| Premium reasoning | $8.00–$25.00 | $40.00–$150.00 | $16.00–$56.25 |
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Questions about future AI token prices are often answered with intuition, isolated vendor announcements or simple extrapolations from today’s pricing tables. We broke the question into the factors that will determine the answer: current provider pricing, historical capability-adjusted price declines, infrastructure roadmaps, inference efficiency, enterprise buying behavior, caching, routing and the changing economics of AI agents.
We treated the cost of reaching a given capability and the list price of the newest frontier model as separate measures. Older capabilities usually become cheaper over time, while providers introduce new premium tiers with stronger reasoning, larger context windows and additional tools. Mixing those two trends produces misleading forecasts.
The forecast uses separate economy, mainstream and frontier categories, with input and output priced independently. Blended prices assume three input tokens for every output token. This ratio is only a common comparison point; actual workloads can look very different.
We evaluated hardware and software efficiency claims alongside the practical limits on retail price declines. Lower compute cost can be passed to customers, absorbed by power and infrastructure constraints, or reinvested in longer reasoning, faster service and larger contexts. We therefore use ranges rather than extending a single historical curve.
Key sources used for this analysis include: OpenAI’s GPT-5.6 announcement, OpenAI API pricing, Anthropic API pricing, Google Gemini API pricing, DeepSeek API pricing, Mistral AI pricing, the Stanford AI Index, Epoch AI research, NVIDIA’s Vera Rubin materials, Artificial Analysis, Berkeley Sky Computing Lab publications, and Microsoft Research publications.
Building or investing in frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inbox