What are you looking for ?
VergeIO
RAIDON

AMD AAI 2026: AMD Launches AMD Helios Rackscale Solution for Frontier AI

New AI infrastructure combines leadership compute, networking and open software, delivering rack-scale performance for at-scale AI

What’s the News? Today at Advancing AI 2026, AMD launched the AMD Helios rackscale solution, a new rackscale AI infrastructure designed for frontier AI, large-scale inference and foundation-model training.AMD Helios combines AMD Instinct MI455X GPUs, 6th Gen AMD EPYC server CPUs, AMD ROCm software and AMD Pensando networking into a unified rack-scale architecture that delivers exaflop-class AI performance, industry-leading memory capacity and bandwidth,1 and exceptional scalability.

Why It Matters? As demand grows for AI factories and frontier AI, customers need infrastructure that scales efficiently from a single rack to large gigawatt-scale AI clusters while maximizing performance, performance per watt and cost per token. AMD Helios delivers an integrated rack-scale building block that simplifies deployment while providing the compute, networking and open software foundation needed for the next generation of AI infrastructure. AMD Helios rackscale AI solution offers the highest token throughput performance at low, medium and high interactivity levels,2 designed with an open architecture from the start. Compared to the leading competitive solution, AMD Helios delivers 15% more peak FP4 performance,3 50% more high-bandwidth memory (HBM) capacity,1 6% more HBM bandwidth1 and 50% more scale-out bandwidth.4 It delivers up to 30% more tokens per dollar than the competitive solution, giving customers more output with every rack they deploy.5

What’s the Role for AMD?
“Frontier AI infrastructure is evolving rapidly, and customers need platforms that can deliver exceptional performance while scaling efficiently as workloads grow. AMD Helios brings together leadership compute, high-performance networking and open software in a unified rack-scale platform. Together, these technologies give customers the flexibility to accelerate large-scale inference, frontier-model training and the next generation of AI infrastructure,” said Vamsi Boppana, SVP, AI, AMD.

What is the AMD Helios Rackscale Solution?
AMD Helios extends the AMD AI portfolio by delivering an open rack-scale architecture designed to support large-scale inference, frontier-model training and fine-tuning from a single rack to multirack clusters. AMD delivers AMD Helios at the same time AI is reshaping modern compute infrastructure, requiring tightly integrated platforms that combine high-performance compute, rack-scale architectures and open software.

  •  Built for Frontier AI: AMD Helios is the AMD rack-scale AI infrastructure platform designed for frontier AI and AI factory deployments. It provides a scalable foundation for high-volume inference, frontier-model training and fine-tuning by bringing together AMD 5th Gen Instinct GPUs, 6th Gen AMD EPYC server CPUs, AMD Pensando networking and AMD ROCm software into a unified platform built on open standards

  • Leadership Rack Performance: AMD Helios delivers exaflop-class AI performance in a unified architecture optimized for large-scale inference and frontier-model development. A single rack provides up to 2.9 exaflops peak FP4, 1.4 exaflops peak FP8, 31 terabytes of HBM4 memory and 1.7 petabytes/second of memory bandwidth

  • Scalable AI: AMD Helios is designed to scale from a single rack to gigawatt-scale AI clusters. Each rack integrates 18 open rack wide-aligned 4-GPU compute trays (72 total GPUs) that provide consistent building blocks for cluster deployment and expansion. At cluster scale, AMD Pensando Vulcano 800 AI NICs provide the high-bandwidth, low-latency connectivity required to support distributed inference, frontier-model training and rapid hyperscale growth

  • Open Infrastructure: AMD Helios combines UALink over Ethernet (UALoE) for high-performance scale-up with standards-based Ethernet aligned with the Ultra Ethernet Consortium for scale-out, enabling customers to build AI infrastructure using open, standards-based technologies. The platform also incorporates defense-in-depth rack-scale security features. AMD ROCm™ software enables high-throughput inference and unified distributed training across rack and cluster-scale deployments

1. Calculations by AMD Performance Labs in June 2026 based on published memory capacity and memory bandwidth specifications of an AMD Helios rackscale solution vs. the published memory capacity and memory bandwidth specifications of a Nvidia Vera Rubin NVL72 rack. System manufacturers may vary configurations, yielding different results
2. Based on modelling by AMD Performance Labs in July 2026, on an AMD Helios rackscale solution to measure token throughput/GPU at high, medium and low interactivity points using Kimi K2 Thinking with an ISL/OSL combination of 32K/8K compared to the published modeled specifications for an Nvidia Vera Rubin NVL72 rack. System manufacturers may vary configurations, yielding different results
3. Based on calculations by AMD Performance Labs in June 2026, to determine the peak theoretical precision performance of an AMD Instinct™ MI455X GPU using the peak Matrix FP16, BF16, INT8 datatypes and peak Open Compute Project MXFP6, MXFP8, FP8 and MXFP4 datatypes vs. published specifications of Nvidia Vera Rubin GPU with the NVFP4 Dense datatype. System manufacturers may vary configurations, yielding different results
4. Calculations by AMD Performance Labs in July 2026, based on published scale-out bandwidth of the AMD rackscale solution vs. published Nvidia Vera Rubin NVL72 Rack specs. System manufacturers may vary configurations, yielding different results
5. Based on AMD Performance Labs estimates as of July 2026, tokens-per-dollar performance was calculated using the Kimi K2 Thinking workload (32K input / 8K output) on an AMD Helios rackscale solution compared to an Nvidia Vera Rubin NVL72 rack. Results reflect estimated aggregate throughput across low, medium, and high-interactivity operating points and hourly pricing projection of system GPUs based on market conditions. System configurations may vary by manufacturer and may produce different results

Read also :

Comments

As many of us saw at different events, AMD Helios is real and during company's Advancing AI 2026 event in San Francisco, CA, AMD launched the AMD Helios rackscale solution, a new rackscale AI infrastructure designed for frontier AI, large-scale inference and foundation-model training. It is a real effort of integration with various developed and acquired elements. Helios combines AMD Instinct MI455X GPUs, 6th Gen AMD EPYC server CPUs, AMD ROCm software and AMD Pensando networking into a unified rack-scale architecture that delivers exaflop-class AI performance, industry-leading memory capacity and bandwidth, and exceptional scalability. The Helios rack itself is pretty impressive with its physical size.

Key specs: a single rack provides up to 2.9 exaflops peak FP4, 1.4 exaflops peak FP8, 31TB of HBM4 memory and 1.7PB/sec of memory bandwidth. Each rack integrates 18 open rack wide-aligned 4-GPU compute trays (72 total GPUs), and AMD Helios is designed to scale from a single rack to gigawatt-scale AI clusters. For networking, Helios combines UALink over Ethernet (UALoE) for scale-up with standards-based Ethernet aligned with the Ultra Ethernet Consortium for scale-out, positioning it as an open alternative to proprietary interconnects, with defense-in-depth rack-scale security features built in.

Now regarding competition, this is where it gets interesting, AMD explicitly benchmarks Helios against Nvidia's next-gen rack, the Vera Rubin NVL72, not the current Blackwell generation. That's a direct shot at Nvidia's roadmap rather than its shipping product. Per AMD's own performance-lab claims:

  • 15% more peak FP4 performance,
  • 50% more HBM capacity,
  • 6% more HBM bandwidth,
  • 50% more scale-out bandwidth versus Vera Rubin NVL72
  • And up to 30% more tokens per dollar than the competitive solution

It illustrates once again that memory has become the battleground. AMD is leaning hardest on HBM capacity/bandwidth advantages, a signal that at rack scale, memory (for KV-cache and long-context inference) is now as strategically important as raw FLOPS.

Cost-per-token, not just peak specs, the pretty famous "30% more tokens per dollar" framing shows AMD competing on economics of inference at scale, the metric hyperscalers actually care about, rather than just headline TFLOPS, an implicit acknowledgment that Nvidia's CUDA/software moat still matters more than pure silicon specs.

On the open ecosystem play, by anchoring on UALink/Ethernet and ROCm, AMD is again betting on open standards vs. Nvidia's NVLink/InfiniBand-centric stack.

All competitive numbers come from AMD based on published specs of Vera Rubin and we wait Nvidia's answers to these especially as Vera Rubin isn't shipping yet either.

Articles_bottom
AIC