Trainium
AWS’s internally designed accelerator family for AI training and related machine-learning workloads.
NVLink Fusion
Nvidia’s technology for connecting custom CPUs and accelerators into Nvidia-compatible rack-scale AI systems.
NVHBM
Nvidia’s custom high-bandwidth memory approach intended to improve bandwidth and power efficiency for AI processors.
AI factory
A large-scale data center system optimized to produce AI output, such as training runs, inference and agentic workloads.
Amazon Press Center
other
AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI
“AWS plans to deploy an additional 2 million NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs in 2027–2028.”
NVIDIA Investor Relations
data
NVIDIA Announces Financial Results for Second Quarter Fiscal 2027
“Revenue of $96.2 billion, up 106% from a year ago; Data Center revenue of $89.0 billion, up 117%.”
NVIDIA Investor Relations
data
NVIDIA Corp. Q2 2027 Earnings Call Corrected Transcript
“AWS is deploying an additional 2 million GPUs starting this quarter through the second quarter of fiscal 2029.”
2M GPUs
AWS and Nvidia plan to deploy 2 million additional Nvidia GPUs across AWS global infrastructure in 2027 and 2028.
$89B data center
Nvidia reported $89.0 billion in data center revenue for fiscal Q2 2027, up 117% from a year earlier.
Trainium tie-in
Amazon’s Annapurna Labs plans to connect future Trainium chips, starting with Trainium4, to Nvidia’s NVLink Fusion architecture.
Amazon Web Services and Nvidia’s latest infrastructure agreement sends a clear near-term signal to cloud infrastructure watchers: the world’s largest cloud buyers may be building their own AI chips, but they still need Nvidia at extraordinary scale.
AWS and Nvidia said on Aug. 26 that they plan to deploy 2 million additional Nvidia GPUs across AWS’s global infrastructure in 2027 and 2028. The companies also said they will expand work on Vera CPUs, networking, open models, government AI factories and robotics systems.1
The commitment follows AWS’s earlier plan to add more than 1 million Nvidia GPUs starting in 2026. The companies said demand had already exceeded those expectations.1
That is the core tension in the deal. Amazon has spent years developing internal silicon, including Trainium for AI training and Graviton for general-purpose compute. Yet AWS is not using that roadmap to step away from Nvidia. Instead, it is buying more Nvidia capacity while integrating future Trainium chips into Nvidia’s rack-scale architecture through NVLink Fusion and NVHBM memory technology.7
For Nvidia investors, the announcement reinforces near-term pricing and platform leverage. For hyperscaler strategists, it suggests custom silicon is becoming a bargaining chip and optimization layer rather than an immediate replacement for Nvidia’s GPU ecosystem.
The AWS order is not coming in a weak demand environment. Nvidia reported $96.2 billion in revenue for its fiscal second quarter ended July 26, up 106% from a year earlier. Data center revenue was $89.0 billion, up 117%. Gross margin was 75.0% on both a GAAP and non-GAAP basis.3
On Nvidia’s earnings call, CFO Colette Kress said AWS is deploying the additional 2 million GPUs beginning in the current quarter and continuing through Nvidia’s fiscal second quarter of 2029. She also said the expansion includes Vera CPUs, some integrated with Rubin GPUs and others deployed as standalone products.4
That timing matters. The public AWS announcement described deployments in 2027 and 2028, while Nvidia’s earnings-call commentary framed the ramp as beginning immediately and running through Q2 FY2029.14
In other words, this is not just a distant capacity reservation. It is part of an active supply plan for AI infrastructure that AWS believes it cannot satisfy with existing capacity or internal accelerators alone.
AWS’s stated reason is customer demand. The company said customers are moving from pilots to production across agentic AI, scientific discovery, enterprise automation and robotics, requiring infrastructure that can keep pace with mission-critical workloads.1
That workload mix favors Nvidia because the company is selling more than GPUs. It is selling a broader platform: accelerators, CPUs, networking, software libraries, models and deployment patterns.
The most important strategic detail may be buried beneath the headline GPU count. AWS and Nvidia are extending work on NVLink Fusion with NVHBM, Nvidia’s custom high-bandwidth memory technology, in collaboration with Amazon’s Annapurna Labs.17
Nvidia said Annapurna will support NVLink Fusion with next-generation Trainium chips starting with Trainium4, allowing Amazon chips and Nvidia GPUs to operate within a common rack-scale architecture.7 AWS said the arrangement would let Trainium access faster, more power-efficient memory and integrate Trainium and GPUs inside a shared rack-scale design.1
That is a nuanced form of competition. AWS still wants more control over silicon economics, supply chains and workload-specific performance. But by aligning Trainium with Nvidia’s interconnect and memory roadmap, AWS is also acknowledging the value of Nvidia’s system architecture.
The custom chip is not being positioned as a clean substitute. It is being made more useful by attaching it to Nvidia’s infrastructure fabric.
This is why the deal is better understood as dependence with optionality. AWS is preserving the option to shift some workloads to Trainium over time, especially where internal chips can deliver favorable price-performance. But for the highest-demand, broadest-compatibility AI workloads, Nvidia remains the default capacity source.
Hyperscalers keep buying Nvidia at scale not simply because Nvidia has the fastest chips in a given generation. They keep buying because Nvidia capacity is fungible across customers, models and workload types.
On the earnings call, Kress described Nvidia compute as fully utilized across the clouds the company serves and emphasized its role across training, inference and agentic workloads.4 Nvidia’s argument is that one platform can support closed models, open models, small models, large models, cloud workloads and edge workloads without forcing customers into separate hardware and software stacks.4
AWS’s announcement points in the same direction. The expanded collaboration includes Nitro and Elastic Fabric Adapter integration for security and networking, Nemotron models on Bedrock and SageMaker, GPU-accelerated data processing for EMR and OpenSearch, and Nvidia’s physical AI platform for Amazon Robotics.1
Nvidia is also pushing further into the CPU layer. The company said AWS has received its first Vera CPU server and Vera Rubin GPU, after similar deliveries to Oracle Cloud Infrastructure and AI labs including Anthropic and OpenAI.8 Vera is designed for agentic AI workloads that rely heavily on orchestration, tool calls, retrieval, sandboxing and data movement — tasks that do not run on GPUs alone.8
For AWS, that means Nvidia is becoming more deeply embedded in the AI factory stack, not less. A cloud provider can negotiate harder when it has internal silicon. But replacing Nvidia requires more than matching a GPU benchmark. It requires matching software compatibility, networking, memory, developer tooling, supply availability and customer confidence.
The longer-term risk to Nvidia is real. If AWS, Google, Microsoft or other hyperscalers can move a growing share of predictable internal workloads onto custom chips, Nvidia may face pressure on pricing, attach rates or gross margins over time.
Nvidia’s own disclosures show why hyperscaler leverage matters. In its latest 10-Q, Nvidia said revenue remains concentrated among a limited number of direct and indirect customers. One direct customer represented 16% of total revenue in the second quarter, while three direct customers represented 16%, 15% and 13% of revenue in the first half of fiscal 2027.6
Large cloud buyers are not just customers. They are among the few entities with enough volume to reshape procurement economics.
The 10-Q also reported that Nvidia’s gross margin rose to 75.0% in the second quarter from 72.4% a year earlier, helped by improved mix from Blackwell Ultra.6 Those margins are precisely what hyperscalers are trying to attack with in-house silicon.
A successful Trainium roadmap could cap Nvidia’s ability to sustain current margin levels indefinitely, particularly for internal AWS workloads where Amazon controls the software stack and can steer demand.
But margin pressure is not the same as displacement. The Aug. 26 announcement shows AWS choosing both paths: more custom silicon and more Nvidia. It is trying to reduce dependence at the margin while still expanding dependence in absolute terms.
For cloud infrastructure watchers, the AWS-Nvidia expansion is a reminder that AI capacity is now constrained by full-stack execution: chips, memory, networking, power, cooling, software, data center construction and customer deployment readiness. Nvidia’s current advantage is that it can package more of that stack than a standalone accelerator supplier.
AWS’s custom chips may become increasingly important in the next phase of AI cloud economics. They can help Amazon tune infrastructure for specific workloads, lower internal costs and pressure Nvidia in negotiations.
Yet the 2 million-GPU expansion says the immediate problem is not excess Nvidia exposure. It is not having enough Nvidia capacity to satisfy customers.
The near-term takeaway is straightforward: hyperscaler custom silicon is a ceiling on Nvidia’s long-term margins, not yet a replacement for Nvidia’s platform. AWS is building alternatives, but its latest commitment shows that the path to competing with Nvidia still runs through Nvidia.
Comments