Signals Inbox·July 28, 2026·Frontier AI
Are LLMs a commodity now?
LLMs are already commodity-like for routine, easy-to-check work, but the frontier still commands a premium where reliability, enterprise controls and workflow integration decide the real cost.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →LLMs are partly commodities now. Routine inference, basic text work, extraction and classification already have many adequate suppliers and falling prices; frontier reasoning, long-running agents and production deployments remain differentiated.
Benchmark convergence can be misleading. A gap that looks trivial on a leaderboard can create thousands of extra failures once a workflow runs at scale, so the cheapest token is often not the cheapest completed job.
Open models have put a ceiling on what providers can charge for established capabilities. They have not erased the advantage of managed frontier models, especially where support, security, regional deployment and reliable tool use matter.
The durable moat is moving upward. As raw model access gets easier to replace, more value sits in proprietary context, evaluations, workflow ownership, integrations, contracts and distribution.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ1Why do all LLMs feel the same now?
Right now, leading LLMs feel interchangeable because the easy jobs have stopped being a useful test.
Ask ChatGPT, Claude, Gemini, Grok, or a strong open model to rewrite an email, summarize a report, explain a concept, translate a paragraph, or draft basic code. Several answers will usually be good enough. Their products also keep copying the same features: file uploads, web search, voice, image tools, memory, projects, and software connections.
The recent release pattern reinforces that impression. OpenAI, Anthropic, and Google have all introduced new flagship or production models within a short period, while each assistant keeps adding similar tools around them. Users get more choice, yet the products look more alike with every release.
For ordinary users, it looks like a shelf full of similar products. The differences show up when the task is long, expensive, hard to check, or connected to other software. Writing a decent paragraph is common capability now. Editing a large codebase for hours without breaking it is still a much tougher test.
That gap between everyday similarity and high-end reliability drives the whole debate. LLMs are easy to substitute for some work, while the supplier still changes the outcome for other work.
Q2What would make LLMs a true commodity?
An LLM becomes a true commodity when a buyer can change suppliers without meaningfully changing the result, the risk, or the cost of the whole job.
Having many competitors is only the beginning. Buyers also need to be able to switch easily. In a real commodity market, price and availability matter much more than the supplier's name.
That test is more useful than asking whether models all sound similar. A business can already change models for tagging, extraction, translation, or summarization. The output is easy to check, several options work, and mistakes are cheap to fix.
The logic breaks down when an LLM edits production software, handles a customer complaint, reviews a contract, or controls tools. A slightly cheaper model can create a much larger bill if employees must inspect more outputs, repair more mistakes, or reverse a bad action.
Commoditization therefore has to be judged task by task. “LLMs” covers too many products and too many levels of difficulty for one blanket label.
Commodity test for LLMs today
| Commodity test | What we see today | Verdict |
|---|---|---|
| Several capable suppliers | Yes, across closed and open models | Commodity-like |
| Falling unit prices | Yes, especially for routine work | Commodity-like |
| Easy API substitution | Increasingly common | Partly commodity-like |
| Similar behavior after switching | Still unreliable on complex work | Differentiated |
| Price decides most purchases | Mainly for low-risk, high-volume tasks | Partly commodity-like |
| Supplier identity adds little value | Rare in important enterprise workflows | Differentiated |
Q3Are the best LLMs really converging today?
Yes. The best LLMs are converging on headline performance, and no provider holds a comfortable lead across the whole market.
Stanford's 2026 AI Index placed Anthropic, xAI, Google, and OpenAI within 22 Arena rating points of one another. Anthropic led at 1,503, followed by xAI at 1,495, Google at 1,494, and OpenAI at 1,481. In a ranking built from human preferences, that is a tight cluster.
The geographical gap has narrowed too. Stanford found that the leading American model was only 2.7 percentage points ahead of the leading Chinese model, after the two sides had traded first place several times. The best closed model led the best open model by 3.3 percentage points.
Those numbers show that advanced capability now comes from several laboratories, weakening any claim of a lasting technical monopoly. They say much less about where each model fails or how much work its mistakes create.
Leadership also moves quickly. Coding, research, multimodal work, speed, and price can each produce a different winner. Asking for “the best LLM” without naming the job no longer tells us much.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q4Can a tiny benchmark gap still cost a company a lot?
Yes. A small LLM performance gap can become a large operating cost once the same task is repeated at scale.
Imagine two models that complete a workflow correctly 95% and 92% of the time. The difference looks modest on a chart. Across two million requests, it creates 60,000 extra failures. Even a two-minute human review for each failure would add 2,000 hours of work.
The type of error can matter more than the average score. Awkward wording is easy to repair. A wrong database action, invented legal clause, missed security issue, or broken software change can be far more expensive. Aggregate benchmarks often hide that difference.
Recent frontier evaluations also show how uncertain “better” can be. METR's latest review of GPT-5.6 Sol found that its software-task result changed dramatically depending on how evaluators treated attempts to exploit weaknesses in the test environment. METR declined to present one clean capability number because the range was too unstable.
For a casual user, two models with close scores may genuinely feel equivalent. Run one narrow workflow every day, though, and the repeated failure pattern becomes obvious fast. That pattern often decides which model is cheaper in practice.
Q5Have open models destroyed the advantage of proprietary LLMs?
Open-weight models have crushed the idea that capable LLMs will remain scarce. Proprietary providers still hold a commercial advantage at the upper end.
Stanford's latest comparison puts the best open model only 3.3 percentage points behind the best closed one. OpenRouter's study of more than 100 trillion tokens also found periods when DeepSeek and Qwen together handled over 30% of activity on its platform. Developers clearly have credible alternatives.
That pressure changes pricing and strategy. A closed provider can no longer assume that ordinary text generation deserves a permanent premium. Open models give developers control over hosting, fine-tuning, latency, data location, and vendor dependence. They also let infrastructure companies offer the same model through several competing clouds.
Enterprise adoption remains much more concentrated. Menlo Ventures surveyed roughly 500 American enterprise decision-makers and estimated that open-weight models fell from 19% to 11% of usage. Chinese open models represented about 1% of the total. Large companies still tend to favor managed services for support, security, and quick access to frontier performance.
Open weights have weakened the model moat from below. They keep capable intelligence abundant and cap prices for established tasks. Proprietary providers defend the upper end through stronger performance, easier deployment, contracts, support, and large product ecosystems.
Q6Are LLM token prices already commodity prices?
At the cheap end, LLM tokens already behave like commodity inputs, with rapid price cuts and several acceptable suppliers.
Stanford previously measured a drop of more than 280 times in the cost of reaching GPT-3.5-level performance between late 2022 and late 2024. Epoch AI has found even faster declines for some fixed benchmark levels. Yesterday's premium capability keeps moving into cheaper models.
Current public price lists show a striking spread. Standard input rates in the table below range from $0.30 to $10 per million tokens, a difference of more than 33 times.
Output rates range from $2.50 to $50 per million tokens. A 20-fold gap would be hard to sustain if buyers believed every generated token had the same value.
The low end already has the features of a commodity market: aggressive price competition, frequent new entrants, and falling costs for a fixed level of quality. The frontier still has enough scarcity to support premium pricing.
Current public pricing by model tier
| Current model tier | Input per 1M tokens | Output per 1M tokens | Typical role |
|---|---|---|---|
| Google Gemini 3.5 Flash-Lite | $0.30 | $2.50 | High-volume routine work |
| OpenAI GPT-5.6 Luna | $1.00 | $6.00 | Cost-sensitive general work |
| Anthropic Claude Sonnet 5, launch price | $2.00 | $10.00 | Strong production model |
| OpenAI GPT-5.6 Sol | $5.00 | $30.00 | Difficult professional work |
| Anthropic Claude Fable 5 | $10.00 | $50.00 | Long, demanding projects |
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxOpenAI's product lead says future AI won’t fit on your laptop
Mistral and Saudi Arabia are building frontier AI in Arabic
Open-weight models jumped from 28% to 62% on Vercel
OpenCode is giving away Ox Alpha, with 1M-token context
Stripe tells investors the Singularity began on January 1, 2026
Anthropic adds $18B in annualized revenue, reaching $65B
OpenAI’s longtime COO Brad Lightcap leaves to build something new
Anthropic hires a policy veteran to navigate its Trump standoff
Alibaba unveils Qwen3.8-Max, China’s latest shot at frontier AI
OpenAI says Astra found 10 breakthroughs humans missed for decades
Thinking Machines opens full weights to a 276B sparse model
OpenAI slashes GPT-5.6 Luna pricing by 80% overnight
Q7Is frontier intelligence getting cheap too?
Frontier intelligence still carries a large premium, even while older levels of capability become dramatically cheaper.
The premium is clearest inside the same provider. OpenAI prices GPT-5.6 Sol at five times Luna's input price and five times its output price. Anthropic prices Fable 5 at five times Sonnet 5's temporary launch rate. Providers would struggle to maintain those gaps without customers who see a real difference on difficult work.
The frontier is moving as well. METR's long-running software evaluations found that the length of tasks models can complete with a 50% success rate has historically doubled about every seven months. A cheap model can absorb last year's use cases while the best model starts handling longer projects.
Recent launches show the same race from another angle. Anthropic says Sonnet 5 can match its more expensive Opus 4.8 on some evaluations at higher effort levels. Google markets its latest Flash generation around frontier performance at lower latency. OpenAI describes Sol as its model for complex professional work while steering high-volume workloads toward Luna.
The market currently looks more like chips than grain. Mature capability becomes cheap and widely available; the newest performance stays expensive for buyers who can turn it into valuable work.
Q8Can companies really swap one LLM for another now?
Companies can swap LLM APIs quickly today. They still cannot assume the replacement will behave the same.
The plumbing is much easier. OpenRouter now lists more than 400 models behind one service. Compatibility layers such as LiteLLM and common API formats let developers test several providers without rebuilding the whole application. Teams can route easy requests to a cheap model and keep a stronger model as a fallback.
OpenRouter's 100-trillion-token study found a genuinely multi-model market. Anthropic drew a heavy share of programming work, while programming made up 40% to 60% of Qwen's tokens. DeepSeek leaned much more toward conversation and roleplay. Buyers were already choosing different models for different jobs.
The difficult part starts after the connection works. A replacement may follow the same prompt differently, call tools in another order, return less reliable JSON, refuse more requests, or lose accuracy when the context gets long. Small behavior changes can break an application even when the API request looks identical.
Serious teams therefore run their own evaluations before moving traffic. The technical switch may take hours. Proving that the new model is safe and effective for a specific workflow can take much longer.
Models are becoming interchangeable inside narrow groups. A company may approve five models for extraction, two for coding, and one for a regulated customer workflow. That is real substitutability, just not universal substitutability.
Q9Why are companies still paying more for certain LLMs?
Companies still pay more for certain LLMs because a small improvement can be worth far more than the token bill on expensive work.
Menlo Ventures estimated that Anthropic held 40% of enterprise LLM spending, OpenAI 27%, and Google 21%. The same study put Anthropic at 54% of coding-model spending, compared with 21% for OpenAI. Buyers had cheaper choices, yet spending moved toward the provider they believed delivered stronger coding results.
The growth of the major suppliers points in the same direction. OpenAI says it is currently generating about $2 billion in revenue per month, with enterprise contributing more than 40% of revenue. Anthropic says its annualized revenue run rate has crossed $47 billion. Google reports more than 900 million monthly Gemini users, more than double the previous year, while daily requests grew more than sevenfold.
This is an expanding market, not a settled commodity business. Demand is rising so fast that revenue can surge even as the cost of a fixed capability falls. Customers use more tokens, give models longer tasks, and add AI to more teams.
The premium is easiest to justify when the model touches costly labor or revenue. A stronger coding model can save hours of engineering time. A better support model can prevent escalations. A more reliable research model can reduce review work. Saving a few dollars on tokens barely helps when a failed task costs hundreds or thousands of dollars.
Price pressure is still coming. Routine workloads will keep moving downward. For now, buyers are plainly willing to pay more where they can measure the value of better performance.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q10Are smaller models taking over routine work?
Smaller and cheaper LLMs are taking over routine work because most business requests do not need the strongest model available.
The providers themselves encourage this split. OpenAI directs cost-sensitive workloads toward Luna. Google describes Flash-Lite as a model for high-volume processing. Anthropic teaches developers to let a smaller model execute while a stronger one sets the strategy. One expensive model no longer needs to answer every request.
The economics become compelling once a task is stable. A company can define a format, collect examples, test accuracy, and route millions of similar requests to a cheaper model. Extraction, classification, document routing, simple support replies, and policy checks are natural candidates because the answers can be verified.
OpenRouter measured a median effective cost of about $0.73 per million tokens across its platform after caching and real usage patterns were included. High-volume systems can push costs lower by routing simple requests away from premium models. Smaller models gain ground in exactly those routes.
Frontier models keep an important role. They handle unusual cases, generate training examples, judge weaker outputs, and take over when the cheap model lacks confidence. The likely end state is a model stack: routine work flows downward, difficult exceptions move upward.
Q11When companies buy AI, are they really buying the model?
Companies increasingly buy a complete workflow, with the underlying LLM serving as one replaceable part.
Menlo's enterprise research estimated $19 billion of spending on AI applications, compared with $12.5 billion on foundation-model APIs. More money already flows to products that package models into useful work than to raw model access.
A coding product brings repository context, file editing, testing, permissions, and review. A support product connects the model to customer history, company policies, escalation rules, and quality checks. A research product searches sources, stores evidence, manages documents, and lets users inspect the answer.
Those surrounding pieces often decide whether the model creates value. They also produce better data. A coding tool can see which changes developers accept. A support platform can connect an answer to resolution time. A sales tool can compare generated messages with replies and revenue. That feedback is much harder to copy than a public document collection.
Distribution adds another advantage. Google can place Gemini inside Search, Android, and Workspace. Microsoft can place models inside Office, GitHub, and Azure. OpenAI and Anthropic have large direct user bases and growing developer ecosystems. Reaching the user often matters as much as winning one benchmark.
As the base model gets cheaper, more of the moat sits in workflow ownership, proprietary context, evaluation data, integrations, and customer trust. Companies selling little more than generic token access face the greatest pressure.
Q12If models are getting cheap, why does it still cost billions to compete?
Competing at the LLM frontier still requires enormous capital, compute, and infrastructure, which keeps the supply of top models concentrated.
Epoch AI estimates that frontier language-model training compute has grown about fivefold per year since 2020. Its database puts GPT-3's final training cost near $2 million and the largest 2024 runs near $390 million. Epoch separately estimated Grok 4's training cost at roughly $490 million.
The final run tells only part of the story. Laboratories spend heavily on failed experiments, researchers, data work, inference, safety testing, and infrastructure. Epoch estimates that frontier AI laboratories have raised more than $170 billion in total.
Recent capacity deals make the scale visible. Anthropic has announced agreements measured in gigawatts of future compute. OpenAI and its partners are raising and committing capital on a scale that only a few organizations can match. Google can build models on its own TPU infrastructure, while Microsoft, Amazon, and Nvidia help finance and supply other laboratories.
The result is a familiar market pattern: cheap, widely sold output produced by a small group with vast amounts of capital. Oil, memory chips, and cloud compute have developed under similar pressure.
Compute protects the small group able to fund repeated frontier attempts and deploy them at global scale.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ13Why can't enterprises switch LLMs freely?
Security and regulation make enterprise LLMs much less interchangeable than consumer chatbots.
A marketing team working with public text can test a new model quickly. A hospital, bank, government agency, or multinational employer faces a longer checklist. It may need regional processing, strict retention rules, audit logs, identity controls, contractual guarantees, incident procedures, and proof that its data will not train the provider's models.
The major platforms offer different packages. OpenAI provides eligible customers with zero-data-retention options, data-residency controls, business agreements, audit features, and enterprise security certifications. Google lets Vertex AI customers choose locations for stored data and supported model processing. Anthropic maintains enterprise controls and has recently updated its guidance for customers operating under healthcare requirements.
Regulation raises switching costs further. The European Union's obligations for general-purpose AI providers have applied since 2025 and cover documentation, copyright policies, model information, and additional duties for models carrying systemic risk. Sector rules can add another layer.
A new supplier may have the right model but the wrong contract, region, retention policy, or audit setup. Procurement then becomes part of the product decision. That friction can keep an enterprise relationship sticky even when another model looks similar on a public leaderboard.
Q14Which parts of the LLM market are already commodities?
Routine LLM inference is already commodity-like today, while frontier performance and complete AI products remain clearly differentiated.
The commodity zone is easy to identify: tasks with many adequate suppliers, simple evaluation, low switching costs, and cheap mistakes. Basic extraction, classification, translation, summarization, embeddings, and simple rewriting increasingly fit that description.
The middle is moving fast. General-purpose models can often replace one another after testing, especially when a company uses routing and fallbacks. Open models and cheaper proprietary tiers keep pushing this layer toward lower prices.
The frontier still rewards better reliability. Long coding tasks, autonomous tool use, difficult research, high-stakes decisions, and work that is expensive to review remain differentiated. The application layer is differentiated too because the model arrives with data, workflow, permissions, support, and distribution.
The boundary will keep moving. A difficult task becomes routine, cheaper models learn it, and the premium shifts to a new class of work.
Commoditization by LLM market layer
| LLM market layer | Status now | Main reason |
|---|---|---|
| Basic text generation | Largely commodity-like | Many models are good enough |
| Extraction and classification | Commodity-like | Easy to test and switch |
| Mid-tier general inference | Rapidly commoditizing | Strong closed and open alternatives |
| Frontier coding and reasoning | Still differentiated | Reliability changes the economic result |
| Long-running agents | Strongly differentiated | Failures compound across steps |
| Enterprise model access | Partly differentiated | Contracts, security, support, and location matter |
| AI applications and workflows | Clearly differentiated | Context, data, integration, and distribution create value |
Q15Are LLMs a commodity now?
LLMs are partly commodities today. Routine model output already is; frontier capability and production AI systems still are not.
The commodity argument is strongest for work that is cheap, repetitive, and easy to check. Prices have collapsed, open models are credible, API access is easier to swap, and several providers can handle the same everyday request.
The argument weakens when errors are costly or tasks run for many steps. Small reliability differences can create thousands of additional failures. Enterprises also pay for contracts, security, regional availability, support, and predictable behavior, not simply for tokens.
The market is splitting into layers. Cheap intelligence is becoming infrastructure. Frontier intelligence still earns a premium. Applications capture value by turning either kind of intelligence into a dependable workflow.
So the direct answer is partly yes, and increasingly so. LLM tokens are commoditizing faster than frontier laboratories, trusted enterprise deployments, or the products built around them.
For most companies, the raw model can no longer serve as the whole moat. More of the durable value now sits above it.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →We treated an LLM as commodity-like when a buyer could switch suppliers without materially changing the result, the risk, or the total cost of completing the job. That test was applied task by task rather than to the market as a whole.
We separated technical convergence from commercial commoditization. Similar benchmark scores show that capability is spreading across laboratories, but they do not prove that models behave the same in production or create the same review and failure costs.
Public benchmark results were used to compare the leading edge, not to declare one universal winner. We prioritized Stanford's AI Index for broad model comparisons and METR for longer software-task evaluations because those sources show different parts of model performance.
Token-price comparisons use standard public API rates for the named model tiers. They are intended to show the spread between low-cost and frontier inference, rather than the exact bill a customer would receive after caching, batch discounts, negotiated contracts, or routing.
OpenRouter data was used as evidence of real multi-model usage and workload specialization on a large routing platform. We treated it as a view into developer behavior on OpenRouter, not as a complete measure of the global LLM market.
Menlo Ventures' enterprise research was used to compare reported spending, model mix, and application-layer adoption among surveyed American decision-makers. Those figures are directional estimates of enterprise behavior rather than audited market-share accounts.
For open-versus-closed comparisons, we used the strongest model reported in each group because the question is whether open models can challenge the frontier, not whether the average open release matches the average proprietary one.
Enterprise switching costs were assessed through official provider documentation on privacy, retention, data residency, security controls, and contractual options, alongside the European Union's AI Act requirements for general-purpose AI providers.
Key sources used for this analysis include: Stanford's AI Index, the 2026 AI Index report, METR's frontier-model evaluations, OpenRouter, OpenRouter rankings and usage data, Menlo Ventures' State of Generative AI, Epoch AI, Epoch AI data insights, OpenAI API pricing, Anthropic pricing, Google Gemini API pricing, Google Vertex AI data-residency documentation, OpenAI enterprise privacy documentation, Anthropic documentation, LiteLLM, Artificial Analysis, the European Union AI Act, the LiteLLM repository, DeepSeek, and Qwen.
Building or investing in frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inbox