AWS will deploy an additional 2 million NVIDIA GPUs across its global infrastructure in 2027 and 2028, after demand already exceeded the more-than-1-million commitment announced at GTC 2026. The chips include Blackwell Ultra, Rubin and Rubin Ultra parts, and the companies are also bringing NVIDIA Vera CPUs to AWS while building secure AI factories for the U.S. government.
The move deepens a 16-year partnership even as Amazon continues to scale its own Trainium accelerators. Matt Garman, CEO of AWS, and Jensen Huang, founder and CEO of NVIDIA, both framed the deal as a direct response to customers racing from pilots into production agentic AI, scientific work, automation and robotics.
The scale of the revision matters as much as the headline count. A plan presented as ample at GTC required a second, larger tranche within months. That pace signals how quickly production agentic and physical AI workloads are consuming reserved capacity once they leave the pilot stage.
The New Order Triples Earlier GPU Plans
At GTC earlier this year AWS said it would add more than 1 million NVIDIA GPUs starting in 2026. Five months later that figure has been overtaken. The fresh commitment covers another 2 million GPUs for 2027-2028, including AI factories, for a combined multi-year footprint that now exceeds 3 million.
NVIDIA’s newsroom release names the specific architectures and notes collaboration on Spectrum networking to link the clusters more efficiently for large-scale training. No dollar value was disclosed, though industry coverage estimates the order in the tens of billions given list prices for high-end GPUs.
| Commitment | Timeline | Scope |
|---|---|---|
| Prior GTC plan | Starting 2026 | More than 1 million NVIDIA GPUs |
| New expansion | 2027-2028 | 2 million additional NVIDIA GPUs (Blackwell Ultra, Rubin, Rubin Ultra) |
| U.S. government AI factories | Part of expansion | 100,000 GPUs on secure AWS infrastructure |
| Combined multi-year | 2026-2028 | More than 3 million GPUs total |
Huang said demand is running ahead of every forecast and that the pair is expanding across the full stack to make agentic and physical AI real at pace only they can deliver.
Tripling the earlier GPU plan in a single follow-on cycle also locks supply windows that rival buyers will find harder to match. Multi-year visibility through 2028 gives AWS scheduling room for power, cooling and networking build-outs that must arrive in step with the silicon.
The inclusion of Blackwell Ultra, Rubin and Rubin Ultra in one order further compresses the usual generation gap. Customers can plan migration paths across those architectures inside a single procurement rather than renegotiating each step.
Sixteen Years From the First Cloud GPU
AWS and NVIDIA launched the world’s first GPU-accelerated cloud instance in 2010. That early CG1 or M2050 hardware was aimed at scientific simulation and graphics. Over the following decade the portfolio grew through K80, V100, A100, H100 and now Blackwell generations, plus UltraClusters that can span thousands of GPUs.
Today AWS claims the widest range of NVIDIA GPU solutions of any cloud provider. The same collaboration produced Project Ceiba, a large GH200-based system hosted on AWS for NVIDIA’s own research, and successive DGX Cloud offerings. The new deal simply extends that pattern into the Rubin era and into CPUs.
The arc from a single instance type to multi-million-GPU factories shows how the partnership absorbed each wave of demand without breaking the operational model. UltraClusters already proved that thousands of GPUs could be scheduled as one fabric. The 2027-2028 build repeats that design at still larger scale.
- 2010 – First GPU-accelerated cloud instance (CG1 or M2050) for scientific simulation and graphics.
- 2010s into early 2020s – Portfolio expands through K80, V100, A100, H100 and Blackwell, with UltraClusters spanning thousands of GPUs.
- Recent collaboration – Project Ceiba on GH200 and successive DGX Cloud offerings hosted on AWS.
- GTC 2026 – Commitment to more than 1 million NVIDIA GPUs starting in 2026.
- Five months later – Additional 2 million GPUs for 2027-2028, Vera CPUs, and secure government AI factories.
Each step widened the customer base from research niches to production agentic fleets. The latest order continues that widening rather than opening a separate track.
Trainium Keeps Growing Beside the GPU Flood
Amazon has spent more than a decade and billions building Annapurna Labs silicon. Trainium and Inferentia were designed to give customers a lower-cost alternative to NVIDIA GPUs and to reduce AWS’s own dependence on a single supplier. Custom-chip revenue run-rates have been reported above $20 billion annually, with large commitments from customers such as Anthropic.
Yet the 2-million-GPU order arrives while AWS is still expanding Trainium. The two paths are no longer framed as either-or.
- Next-generation Trainium chips will use NVIDIA NVLink Fusion interconnects so GPUs and Trainium can sit in the same rack-scale systems.
- Annapurna Labs is working with NVIDIA’s new custom high-bandwidth memory (NVHBM) for faster, more efficient memory on Trainium.
- Customers retain the choice of NVIDIA GPUs, Trainium, or both on the same Nitro-secured, EFA-networked fabric.
- Vera CPUs will arrive as both host processors for accelerated systems and standalone options for agentic workloads that still need heavy CPU orchestration.
The irony is plain. Years of work to escape lock-in now sit inside a deeper technical marriage. Total AI compute demand is simply rising faster than any single hyperscaler’s substitution roadmap. Parallel scaling of custom silicon and merchant GPUs has become the default, not a temporary phase. Similar multi-year capacity talks, including recent Nvidia guarantee talks for OpenAI capacity, show the same pattern across the industry.
NVLink Fusion and NVHBM turn what could have been competing roadmaps into a shared rack design. A customer can place Trainium and NVIDIA GPUs in one system, keep the Nitro security boundary, and move traffic over the same EFA fabric. That arrangement preserves the cost lever of custom silicon without forcing a full cutover away from merchant GPUs.
Anthropic-scale commitments on Trainium already showed that large buyers will take custom chips when price and supply line up. The new GPU tranche shows those same buyers still need merchant capacity for other model families and training shapes. Both lines grow because the workload mix itself is broadening.
Vera CPUs Target the Agentic Loop
Agentic AI spends large stretches of time outside the GPU. Code execution, tool calls, sandboxing, data pipelines and orchestration all land on the CPU. NVIDIA designed Vera for exactly that loop.
Each Vera chip carries 88 custom Olympus cores, up to 1.2 TB/s of memory bandwidth and claimed 1.8x sustained per-core performance versus conventional x86 designs on agentic workloads. It can serve as the host CPU for Vera Rubin NVL72 systems or stand alone inside AI factories. AWS will offer both configurations, giving customers another dial to turn when they size agent fleets.
Early deliveries of Vera systems have already gone to top labs. The AWS arrival extends that footprint into the largest public cloud.
The agentic loop is bursty. GPUs handle the heavy inference or training steps, then the CPU must marshal tools, enforce sandboxes and feed the next prompt. A weak host processor leaves accelerators idle. Vera’s bandwidth and per-core claims aim at that handoff delay.
Offering Vera both as the NVL72 host and as a standalone CPU lets operators match silicon to the duty cycle of each fleet. Orchestration-heavy agents can run on Vera alone. Training-heavy jobs can keep Vera next to Rubin GPUs inside the same tightly coupled system.
Secure Factories for Federal Workloads
One hundred thousand of the new GPUs are earmarked for AI factories built for the U.S. government. The stack will run on AWS infrastructure cleared for workloads classified at Impact Level 6 and above, among the highest civilian security levels. That matches earlier Pentagon moves to put commercial AI models and hardware onto classified networks for operational use.
Garman positioned the federal piece as part of a broader offer to frontier labs, enterprises and governments that want both choice and confidence the pieces work together. Huang called the overall expansion a growth engine already outrunning forecasts.
One hundred thousand GPUs is a small slice of the multi-million total, yet it is large enough to support sustained classified training and inference rather than isolated experiments. Impact Level 6 clearance on the same Nitro and EFA foundation used for commercial regions reduces the need for a wholly separate software stack.
The federal factories also test the full-stack story under stricter controls. Spectrum networking, Vera hosts and Rubin-class GPUs must operate inside the cleared boundary without losing the efficiency gains promised for open regions. Success there strengthens the same architecture for regulated industries outside government.
What Customers Can Already Run
The forward commitments rest on integrations that are live today. All NVIDIA GPU and Trainium EC2 instances sit on the AWS Nitro System and connect through Elastic Fabric Adapter. NVIDIA’s Nemotron open models are available as managed serverless options on Amazon Bedrock and for self-managed fine-tuning on SageMaker.
| Capability | Claimed Gain | Where It Runs |
|---|---|---|
| GPU-accelerated Spark on EMR with cuDF | Up to 3.7x faster processing, 30% better price-performance | Amazon EMR |
| GPU vector indexing | Up to 9x faster indexing at one-quarter the cost | Amazon OpenSearch Service |
| Physical AI stack | Simulation, synthetic data, real-to-sim validation | Amazon Robotics with Jetson, Omniverse, Isaac |
Amazon Robotics is folding the full NVIDIA physical AI platform into warehouse automation. That closes the loop from cloud training to real robots on the floor. Edge and portable agent work is also advancing; recent local Nvidia-powered agent hardware shows the same stack moving outside the data center.
Live services matter because the 2027-2028 GPUs will land on a platform customers already know how to operate. EMR, OpenSearch and Bedrock paths do not need to be redesigned around the new architectures. The same Nitro and EFA layer will carry traffic for Blackwell Ultra and Rubin clusters when they arrive.
Physical AI adds a second payoff path. Models trained in the cloud can be validated in Omniverse, pushed to Jetson-class edge hardware and run on warehouse robots without a separate vendor chain. That continuity is part of what Garman and Huang mean by making agentic and physical AI real at pace.
Spectrum Networking Stitches Clusters at Scale
Large training jobs fail as often on the network as on the GPU. Spectrum networking is the piece NVIDIA and AWS are aligning so that multi-thousand-GPU domains stay efficient as the footprint climbs past three million accelerators over the multi-year window.
EFA already provides the customer-facing fabric for GPU and Trainium instances on Nitro. Spectrum extends the same goal deeper into the cluster spine: higher effective bandwidth between racks, tighter latency control for collective operations, and fewer stranded accelerators when jobs span halls rather than single racks.
Without that layer, adding another 2 million GPUs would raise raw peak flops while leaving a larger share of cycles waiting on all-reduce and parameter exchange. The collaboration treats networking as a first-order capacity product, not a follow-on upgrade.
- Spectrum links training clusters for large-scale jobs called out in the expansion.
- EFA remains the instance-level fabric customers already use across GPU and Trainium fleets.
- NVLink Fusion ties GPUs and Trainium inside shared rack-scale systems.
- Together the three layers keep host CPUs, custom accelerators and merchant GPUs on one operational design.
That design is what lets AWS quote a combined multi-year footprint above 3 million GPUs without promising a clean-sheet network rebuild for every generation.
Dual Paths Now Shape Buyer Choices
Buyers reading the order face a clearer menu than the either-or framing of earlier custom-chip generations. NVIDIA GPUs, Trainium, Vera hosts and Vera standalone CPUs can share security and networking assumptions even when the silicon differs.
The practical effect is portfolio scheduling. A lab can keep frontier training on Rubin-class GPUs, place cost-sensitive fine-tunes on Trainium, and run tool-heavy agents on Vera, all inside the same account boundary. Anthropic-scale custom-chip commitments and the new merchant GPU tranche can coexist in one capacity plan.
| Option | Primary role in the deal | How it joins the fabric |
|---|---|---|
| NVIDIA GPUs (Blackwell Ultra, Rubin, Rubin Ultra) | Core training and high-end inference capacity | Nitro, EFA, Spectrum, NVLink Fusion racks |
| Trainium | Lower-cost alternative and supplier diversity | NVLink Fusion, NVHBM work, same Nitro and EFA fabric |
| Vera CPU | Agentic orchestration and NVL72 host duties | Host for Vera Rubin systems or standalone in AI factories |
| Government AI factories | Classified workloads at Impact Level 6 and above | 100,000 GPUs on cleared AWS infrastructure |
Industry capacity talks, including the OpenAI-related guarantee discussions already noted, follow a similar dual path: secure merchant supply while still funding alternatives. AWS is executing that pattern in public with named architectures and a dated delivery window.
Choice does not remove scarcity. Power, cooling and memory still constrain how fast any of these options can be lit. It does change how customers allocate the scarcity they can reserve through 2028.
Supply Locked While Capex Climbs
The 2027-2028 window gives both companies multi-year visibility that most GPU contracts lack. NVIDIA reported $96.2 billion in second-quarter revenue, with data-center sales at $89 billion, up 117 percent year over year, and guided $108 billion for the following quarter. Analysts have lifted hyperscaler capital-expenditure forecasts into the $1.2 trillion range for 2027.
On X, the sharpest reactions focused less on the headline number and more on the timeline and the coexistence of Trainium with another massive merchant order. One recurring observation was that Amazon is not replacing NVIDIA; it is buying more of both because the overall pie is still expanding. Another noted that Vera now sells into the CPU socket Amazon had largely claimed with Graviton.
Power, cooling, memory and networking costs will rise alongside the GPUs. Whether the resulting AI revenue covers the capital remains the open test for every hyperscaler. For now the order books are full through 2028, and AWS and NVIDIA are filling them together.
The locked supply window turns those cost pressures into a shared planning problem. AWS must bring facilities online on the same calendar as GPU and Vera shipments. NVIDIA must keep Blackwell Ultra and Rubin lines fed while Spectrum and NVLink Fusion land in the same racks. The 16-year working relationship is what makes that joint calendar credible to customers already moving pilots into production.
Hyderabad Gold Dip Freezes Buying as Duty Talk Hits Wedding Season
India Pushes Japan CEPA Review After 15 Years of One-Way Gains
HDFC Bank’s US Class Action Turns Deposit Scheme Into Securities Reckoning
Gold Holds Above $4600 as Treasury Buybacks Fuel Debasement Bid
Scotland’s Prevention Push Leaves Community Groups Squeezed
Nvidia’s 70% Growth Call Carries Its Own Margin Squeeze