Signals Inbox·July 25, 2026·AI Chips
Is Amazon replacing Nvidia?
Amazon is replacing Nvidia across a growing slice of AWS workloads, but not across AWS as a whole—and nowhere near across the wider AI market. Trainium is now a hyperscale platform; Nvidia remains the default wherever portability, mature software and broad distribution matter most.
We track AI chips daily. Want the market signals in your inbox?
Send me the signals →Amazon is replacing Nvidia in selected AWS workloads, but Nvidia remains the dominant AI platform inside AWS and across the wider market.
Trainium’s threat does not depend on beating Nvidia in every benchmark. Amazon can create billions of dollars of value simply by directing more AWS-native work to hardware it designs, operates and rents itself.
The million-chip deployments and gigawatt commitments prove Trainium has reached hyperscale. They do not prove neutral market preference: much of the visible demand comes through investment, cloud-distribution and product partnerships with Anthropic and OpenAI.
Nvidia can keep growing rapidly while losing part of Amazon’s accelerator budget. The market is expanding fast enough for both things to happen at once, which is why revenue growth alone will not reveal the shift immediately.
Inference is likely to be Trainium’s biggest opening. Repetitive workloads reward the cost, power and system-level optimizations Amazon controls, while fast-changing frontier research still favors Nvidia’s broader software ecosystem.
Interested in AI chips?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ1Why does Amazon look like a serious Nvidia threat now?
Amazon looks more dangerous to Nvidia today because Trainium has crossed from a promising internal chip into a platform backed by million-chip deployments, multi-gigawatt contracts and a fast-growing AWS silicon business.
Amazon’s latest quarterly report says its chip business, which includes Trainium, Graviton and Nitro, has passed a $20 billion annual revenue run rate and is growing at a triple-digit percentage. Over the previous 12 months, AWS brought more than 2.1 million AI accelerators into its infrastructure, with Trainium accounting for more than half.
The customer commitments are harder to dismiss. Anthropic has reserved up to five gigawatts of AWS capacity, while OpenAI has committed to roughly two gigawatts of Trainium3 and Trainium4. Those agreements sit inside much broader commercial relationships with Amazon, so they are not neutral chip contests. Still, frontier AI companies do not reserve gigawatts of hardware they think cannot run important production workloads.
A year ago, the strongest Trainium argument was that Amazon needed a cheaper alternative to Nvidia. Now Amazon can point to scale, real customers, a fourth-generation roadmap and software that is improving quickly. Nvidia still leads comfortably, but Trainium is no longer an AWS side project.
Q2What would “replacing Nvidia” actually mean?
Amazon can replace Nvidia for part of AWS’s own AI workload. Replacing Nvidia across the wider market is a much bigger claim, and it is currently false.
There are three useful tests. The first is whether Amazon can use Trainium instead of buying an Nvidia GPU for a specific AWS workload. That is already happening. The second is whether ordinary AWS customers actively move existing work from Nvidia to Trainium. We have some examples, although Amazon does not publish enough usage data to show that this has become normal. The third is whether Trainium can challenge Nvidia across clouds, private data centers and independent AI providers. Amazon has only just started exploring that wider market.
What “replacing Nvidia” would mean
| Meaning of “replace Nvidia” | Current answer | What we can actually verify |
|---|---|---|
| Reduce AWS purchases of Nvidia GPUs | Partly | Trainium handles large internal and customer workloads. |
| Move AWS customers from Nvidia to Trainium | Selectively | Several customers report migrations or new Trainium deployments. |
| Become the default AI accelerator inside AWS | Unproven | AWS does not disclose accelerator hours or revenue by chip platform. |
| Replace Nvidia across the global AI market | No | Nvidia remains available through every major cloud and a large private-infrastructure market. |
Q3Has Trainium reached Nvidia-like scale inside AWS?
Yes, Trainium is already operating at hyperscale inside AWS, although chip counts do not tell us which platform supplies more useful compute.
AWS says more than half of the 2.1 million AI chips it landed over 12 months were Trainium. That puts the minimum above 1.05 million chips. During the same reporting cycle, Amazon announced that more than one million Nvidia GPUs would start deploying in 2026. The two figures sit in the same broad range, but a Trainium2 chip and a Blackwell-class GPU are not interchangeable units of performance, memory, power or cost.
Trainium3 shows how far the system has developed. One Trn3 UltraServer can connect 144 chips and provide 20.7 terabytes of HBM3e memory with 706 terabytes per second of aggregate memory bandwidth. AWS says its UltraClusters can scale to hundreds of thousands of Trainium3 chips. Those are frontier-scale systems, even if the performance claims still come mainly from Amazon.
Trainium and Nvidia scale inside AWS
| AWS scale indicator | Reported level | What it tells us |
|---|---|---|
| AI chips added over 12 months | More than 2.1 million | AWS is expanding accelerator capacity at exceptional speed. |
| Trainium share | More than half | Trainium has passed one million deployed chips on Amazon’s count. |
| Nvidia deployment announced | More than one million GPUs | AWS is expanding Nvidia and Trainium together. |
| Maximum Trainium3 UltraServer size | 144 chips | Amazon offers complete high-density AI systems through AWS. |
We track AI chips daily. Want the market signals in your inbox?
Send me the signals →Q4Did Anthropic and OpenAI prove Trainium can handle frontier AI?
Mostly. Anthropic has proved Trainium can train and serve a frontier model, while OpenAI has shown it is willing to reserve the platform for future production workloads.
Anthropic currently uses more than one million Trainium2 chips for Claude training and inference through Project Rainier. Its expanded agreement covers up to five gigawatts of capacity and more than $100 billion of AWS technology spending over ten years. Anthropic also says more than 100,000 customers run Claude through Amazon Bedrock.
OpenAI provides a different kind of evidence. Its original $38 billion AWS agreement centered on hundreds of thousands of Nvidia GPUs, which shows where mature workloads could run immediately. The later expansion added roughly two gigawatts of Trainium3 and Trainium4 for Frontier, stateful runtimes and other advanced services. Amazon is also investing heavily in OpenAI, so price, distribution and financing all shaped the deal.
Neither company has chosen a single-chip strategy. Anthropic also uses Google TPUs and Nvidia GPUs, while OpenAI spreads work across Nvidia, AMD, Trainium, Cerebras and its planned Broadcom chip. Trainium is proven enough for frontier work. That does not make it the best platform for every frontier model.
Q5Is Trainium really cheaper than Nvidia GPUs?
For large, stable workloads that AWS helps optimize, Trainium can cut costs sharply. For smaller teams, the migration work can swallow much of the saving.
The strongest evidence comes from production case studies rather than Amazon’s headline instance comparisons. Splash Music says Trainium reduced its model-training costs by 54% and training time by 50%. Karakuri reports more than 50% lower LLM training costs. Decart says early tests produced four times more video frames, twice the cost efficiency and latency falling from 40 milliseconds to 10 milliseconds compared with leading GPUs.
These results cover music generation, Japanese-language models and real-time video, so the advantage is not confined to one workload. They also share an important condition: AWS worked closely with the customer on the migration and optimization. That support is realistic for a strategic account. A smaller team with an unusual model may get a rather different experience.
Amazon has a basic economic advantage. It designs the chip, operates the data center and sells the cloud service, so it avoids paying Nvidia’s hardware margin before adding its own. Customers still have to count engineering time, retraining, debugging and deeper dependence on AWS. Trainium can be much cheaper once the whole workload fits.
Q6Do Amazon’s benchmarks prove Trainium beats Nvidia?
No. Amazon has shown strong wins on selected workloads, but it has not shown that Trainium beats Nvidia across a broad, independent test set.
The latest MLPerf Training round included 95 systems using 13 different accelerators. Nvidia systems appeared across many submissions, alongside AMD, Google and several cloud providers. Amazon did not submit a comparable Trainium result, even though the new tests included large and small mixture-of-experts models that should suit Trainium3’s design.
Independent research gives Trainium more credibility than Amazon’s marketing alone. One published project trained 7-billion- and 70-billion-parameter models over 1.8 trillion tokens using 4,096 Trainium accelerators and reached model quality comparable with similar GPU- and TPU-trained models. Another research team improved Trainium inference by an average of 1.66 times after replacing parts of AWS’s matrix-multiplication implementation.
So Trainium can do serious work, while its software still leaves performance on the table. Until Amazon participates consistently in broad public benchmarks, claims that Trainium generally beats Nvidia are too strong.
Interested in AI chips?We can send you all the signals
Send me the signals → Delivered straight to your inboxOpenAI’s Jalapeño beats Nvidia Blackwell on speed and efficiency
Nvidia is eyeing Korea’s $2.3B challenger in AI inference
Cambricon just gave 124 engineers stock worth $828,000 each
Nvidia’s $20B Groq is now entering full production
SK Hynix buys back $29B after shares halve
Nvidia raises AI server prices over 15% starting early 2027
Micron is building a $50 billion chip city inside Boise
Etched ships its first cluster to Jane Street, raises $700M
Groq raises $350M as its valuation falls to $3.5B
SpaceX and Tesla are building a $16.8B gas-powered chip fab
AMD is acquiring Taalas to hardwire AI models into silicon
Huawei targets 1.4nm-equivalent chips by 2031 without EUV
Q7Is AWS Neuron now a credible alternative to CUDA?
Neuron is credible enough for serious AWS projects today, but CUDA still gives developers far more choice, maturity and portability.
AWS now supports PyTorch, JAX, Hugging Face, vLLM, PyTorch Lightning and several managed services around Trainium. A recent Neuron release moved its low-level Kernel Interface and profiling tools out of beta, added a CPU simulator for local debugging and expanded the library of ready-made kernels. These are useful fixes to real developer complaints, not just another layer of product language.
The gap with CUDA remains large. CUDA code and skills travel across AWS, Microsoft Azure, Google Cloud, Oracle, specialist GPU clouds, private servers and workstations. Deep Trainium optimization leads back to AWS infrastructure, even when the original PyTorch model required few changes.
For a company already committed to AWS, Neuron may now be good enough to justify the lower compute bill. A company that values portability or uses many specialized libraries will often keep paying more for Nvidia because the surrounding software saves time.
Q8Are ordinary customers choosing Trainium now?
Some are, but the public evidence still leans heavily toward a few customers receiving unusually close AWS support.
AWS names Anthropic, OpenAI, Databricks, Uber, Poolside, Decart, Karakuri, Ricoh, Splash Music and others as Trainium users or partners. The list covers coding, language models, music, video, transport and enterprise software. AWS has also added Trainium support to more standard services, including managed container infrastructure, which should make adoption less of a special engineering project.
The missing numbers tell us more than another customer logo. Amazon does not disclose how many companies use Trainium in production, how many moved from Nvidia, what share renewed after a trial or how much Trainium revenue comes from customers without investment and distribution ties to Amazon.
Q9Why is Amazon spending so much on its own AI chips?
Amazon wants Trainium because buying every AI accelerator from Nvidia would leave too much margin, supply control and product strategy in someone else’s hands.
Nvidia’s latest quarterly gross margin was about 75%. That margin pays for a full hardware and software platform, but it also shows how much value Nvidia captures before AWS rents the machines to customers. At the scale Amazon is now building, even a modest reduction in accelerator cost can save billions.
Amazon’s latest figures make the pressure visible. Trailing-12-month purchases of property and equipment, net of sales and incentives, reached roughly $147.3 billion, up 67% from a year earlier. Free cash flow fell to about $1.2 billion, with Amazon attributing the drop mainly to a $59.3 billion increase in capital spending driven by AI.
Trainium gives AWS another way to manage that bill. Amazon can tune the chip for expected workloads, decide how much capacity to build and keep more of the margin when customers rent it. The chip does not have to win every benchmark for the strategy to pay off.
We track AI chips daily. Want the market signals in your inbox?
Send me the signals →Q10Will Trainium take Nvidia share first in inference or training?
Inference is the easier path because repetitive, high-volume workloads reward the cost and power optimizations Amazon controls.
Frontier training changes quickly. New model architectures, precision formats and communication patterns can appear between chip generations. Nvidia’s programmable GPUs and mature libraries make it easier for researchers to change direction without waiting for a compiler or kernel update.
Inference becomes more predictable after a model is chosen. Amazon can optimize memory use, batching, speculative decoding and networking around billions of repeated requests. Trainium3 is built around that opportunity: AWS says it can deliver more than five times as many output tokens per megawatt as Trainium2 at similar user latency on Bedrock.
Trainium will keep handling some large training jobs, especially where AWS works directly with the model developer. Broader adoption is more likely to come from serving models, post-training and other workloads where a small saving per token becomes enormous at scale.
Q11Can Trainium compete outside AWS?
Not at Nvidia’s scale. Amazon has only just started trying to break out of the AWS-only box.
Amazon AI chief Peter DeSantis recently confirmed that the company is holding early discussions with organizations interested in using Trainium in their own data centers. No buyers have been named, and Amazon describes the conversations as exploratory. Strategically interesting, commercially unproven.
AWS AI Factories already offer a halfway step. A customer provides the data center and power, while AWS deploys and manages dedicated infrastructure using Trainium accelerators, Nvidia GPUs or both. The infrastructure is local, but Amazon still owns the service relationship and operating layer.
A successful direct-hardware business would change the comparison because Trainium would no longer depend entirely on AWS cloud demand. For now, Nvidia keeps the decisive distribution advantage: its platform is sold through clouds, server makers, independent data centers and enterprise systems worldwide.
Q12Is Nvidia losing its position inside AWS?
No. AWS is adding Trainium quickly while ordering Nvidia at a scale that still looks enormous.
Amazon plans to deploy more than one million Nvidia GPUs beginning in 2026 and continues rolling out Blackwell and Blackwell Ultra systems. Its initial OpenAI infrastructure agreement relied on hundreds of thousands of Nvidia GPUs. AWS and Nvidia have also expanded their work on networking, inference software and future systems, including Nvidia interconnect technology for Trainium4.
The relationship is useful to both sides. Nvidia brings customers to AWS because many developers already depend on CUDA. Trainium gives Amazon a cheaper and more controllable option for workloads that can be optimized around its own stack. AWS would hurt itself by forcing every customer onto Trainium before the software and demand were ready.
This is a change in the mix, not an Nvidia exit. Trainium can capture more of each new AWS data-center build even while Amazon keeps buying more Nvidia hardware in absolute terms.
Interested in AI chips?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ13Does Nvidia’s revenue show any sign of replacement?
No. Nvidia’s latest results look like a company absorbing more demand than the market can currently supply, rather than one being pushed out by Amazon.
Nvidia’s data-center revenue rose from $39.1 billion to $75.2 billion across five reported quarters, an increase of 92% from the first to the last. The latest quarter alone grew 21% sequentially. Custom chips from Amazon and Google expanded during the same period, yet Nvidia added more than $36 billion of quarterly data-center revenue.
The market is growing fast enough for Amazon to take some share without shrinking Nvidia’s business. A real replacement would eventually appear through weaker Nvidia growth, pricing pressure, underused inventory or falling customer commitments. None of those patterns is visible in the latest financial results.
Nvidia data-center revenue progression
| Nvidia reported quarter | Data-center revenue | Sequential change |
|---|---|---|
| Q1 fiscal 2026 | $39.1 billion | — |
| Q2 fiscal 2026 | $41.1 billion | +5% |
| Q3 fiscal 2026 | $51.2 billion | +25% |
| Q4 fiscal 2026 | $62.3 billion | +22% |
| Q1 fiscal 2027 | $75.2 billion | +21% |
Q14Can Amazon hurt Nvidia without overtaking it?
Yes, and this is the most likely outcome. Amazon can lower Nvidia’s share of AWS spending, gain bargaining power and keep more cloud margin.
The damage would first appear inside new capacity decisions. When AWS builds another AI cluster, it can reserve Nvidia for workloads that need CUDA or maximum flexibility and send more predictable jobs to Trainium. Nvidia may continue selling more GPUs every year while receiving a smaller percentage of Amazon’s total accelerator budget.
Trainium also changes negotiations. Amazon has a credible fallback when Nvidia supply is tight or prices are high. AWS can price its own chips aggressively because it earns money from storage, networking, Bedrock and long-term cloud contracts around the accelerator.
That would be financially meaningful even if Trainium never becomes the larger global platform. Nvidia would remain the industry leader, while Amazon would stop being a customer with no practical alternative.
Q15What could stop Trainium from becoming a major Nvidia alternative?
Software friction and customer concentration are the two biggest risks. Chip performance now looks less uncertain than broad, repeatable adoption.
The first problem appears when a workload moves beyond standard PyTorch support. Teams may need Neuron-specific kernels, compiler tuning and debugging skills to reach the savings promised by AWS. Large partners can work directly with Annapurna Labs. A smaller company may decide that Nvidia costs more but wastes fewer engineering weeks.
The second problem is concentration. Anthropic and OpenAI account for a huge share of the visible future demand, and both agreements include investment, cloud distribution and product cooperation. Amazon still needs many customers that choose Trainium mainly because the product is better for their workload.
Hardware timing adds another risk. Amazon must choose memory, interconnect and precision features years before it knows which model designs will dominate. Nvidia faces the same uncertainty, but its larger developer ecosystem can adapt software around new workloads faster. Trainium has become credible. Turning that credibility into a durable platform is the harder part.
We track AI chips daily. Want the market signals in your inbox?
Send me the signals →Q16What would prove Amazon had replaced Nvidia inside AWS?
We would call it replacement only after Trainium becomes the default for normal AWS customers and handles most AI accelerator usage, rather than relying on a few giant contracts.
The cleanest evidence would be a majority of AWS training and inference hours, accelerator-service revenue or power consumption running on Trainium for several consecutive periods. Amazon currently publishes chip counts, customer commitments and broad silicon revenue, but none of those measures tells us which platform carries most actual AI work.
We would also want to see repeated migrations from Nvidia, renewals after the first contract and independent cost comparisons that include engineering time. Public benchmarks should show Trainium performing well across several model families instead of a few carefully optimized cases.
Amazon does not need Trainium to dominate every cloud before it can replace Nvidia within AWS. It does need evidence that customers choose it without a special partnership, a large investment deal or direct help from Amazon’s chip engineers.
Q17Is Amazon replacing Nvidia?
Amazon is replacing Nvidia in a slice of AWS workloads, but Nvidia remains the industry’s default AI platform.
Trainium runs at genuine hyperscale, supports frontier AI work and can deliver large savings for workloads that fit AWS’s stack. Amazon has also built enough leverage to decide that some new AI capacity no longer needs an Nvidia GPU.
Nvidia still has the stronger position. Its software travels across clouds and private infrastructure, its public benchmark record is broader, and its data-center revenue is climbing at an extraordinary rate. AWS itself continues making huge Nvidia commitments while expanding Trainium.
The likely result is a mixed market. Amazon will move more AWS-native training, inference and post-training onto its own chips, reducing Nvidia’s share and weakening its negotiating power inside the world’s largest cloud platform. Nvidia will keep the workloads that value flexibility, portability and the widest software ecosystem.
Amazon is becoming one of Nvidia’s most dangerous customers and competitors. It is nowhere close to replacing Nvidia across the AI industry.
We track AI chips daily. Want the market signals in your inbox?
Send me the signals →This analysis tests whether Amazon is replacing Nvidia by looking at the places where replacement would actually become visible: infrastructure scale, customer adoption, frontier-model usage, workload economics, public benchmark evidence, software maturity, distribution outside AWS, Amazon’s continuing Nvidia purchases and Nvidia’s own financial performance.
We prioritized the freshest material evidence available, including company filings, reported infrastructure commitments, chip deployments, production case studies, technical documentation, independent research and standardized benchmark results. No single contract, benchmark or headline determined the conclusion.
Chip deployments were used to measure scale, but we did not treat one Trainium chip as equivalent to one Nvidia GPU. The platforms differ in performance, memory, power use, system design and cost, so raw unit counts are useful only as broad infrastructure indicators.
Gigawatt commitments were treated as evidence of serious intended usage, not as neutral chip comparisons. The Anthropic and OpenAI agreements include investment, cloud distribution and wider product cooperation, all of which can influence hardware selection.
Customer case studies were used to establish that Trainium can deliver meaningful savings on selected production workloads. We did not assume that those results transfer automatically to smaller customers, different models or teams without direct AWS engineering support.
We used public benchmarks and independent research as a stronger test of broad performance claims. Trainium’s limited presence in standardized benchmark submissions means the evidence supports technical credibility, but not a general claim that it outperforms Nvidia.
We separated absolute growth from competitive displacement. Amazon can move a larger share of new AWS workloads to Trainium while continuing to buy more Nvidia systems in absolute terms. Nvidia can therefore keep growing quickly while losing part of Amazon’s accelerator budget.
We reserved the word “replacement” for broad, sustained and repeatable substitution. A credible alternative, several large deployments or lower costs on selected workloads do not by themselves establish that Trainium has become the default platform inside AWS or across the wider market.
Key sources used for this analysis include: Amazon’s first-quarter results on chip scale, infrastructure spending and Nvidia deployment plans, Amazon’s shareholder letter on custom silicon and Bedrock economics, AWS’s Trainium product documentation, AWS’s Trainium3 UltraServer announcement, Anthropic’s account of Project Rainier and its expanded AWS commitment, OpenAI’s Trainium partnership with Amazon, and the original AWS–OpenAI infrastructure agreement.
Additional sources include: AWS’s Splash Music case study, AWS’s Karakuri case study, AWS’s Trainium customer evidence, the AWS Neuron 2.31 release notes, the HLAT Trainium research paper, MLPerf Training v6.0 results, AWS AI Factories documentation, and Nvidia’s first-quarter fiscal 2027 results.
Building or investing in AI chips?We can send you all the signals
Send me the signals → Delivered straight to your inbox