The takeaway
AMD Helios, built on open standards with 72 GPUs, 31 TB HBM4, and 260 TB/s scale-up bandwidth, signals that the AI hardware market is fragmenting from a single-supplier dynamic into a competitive multi-vendor landscape — good news for builders facing rising compute costs.
Why it matters for builders
AMD Helios gives AI builders a credible alternative to NVIDIA with open Ethernet-based networking, 31 TB HBM4 per rack for large-model inference without sharding, and ROCm software with native PyTorch/JAX support. Lower inference costs from hardware competition expand the market for AI applications.
AMD Helios Rackscale Platform Reshapes AI Hardware Competition
AMD has drawn a line in the silicon. On August 4, the company published the most detailed technical disclosure yet of its next-generation AI infrastructure stack: the CDNA 5 architecture powering the Instinct MI455X GPU, packaged inside Helios, a 72-GPU rack-scale platform that treats the entire rack — not the server — as the fundamental unit of AI compute.
The announcement isn't a paper launch. AMD confirmed that Helios is in production today, with 6th Gen EPYC "Venice" CPUs, 5th Gen Instinct MI455X GPUs, AMD Pensando networking, and the ROCm software stack all shipping as a unified system. The numbers are staggering: 40 petaflops of FP4 compute per GPU, 432 GB of HBM4 memory with 23.3 TB/s bandwidth per GPU, 31 TB of total HBM4 per rack, and 260 TB/s of scale-up bandwidth connecting all 72 GPUs into a single domain.

What Happened
The CDNA 5 architecture represents AMD's most aggressive leap in AI silicon to date. It introduces a 4-bit Tensor Lookup Table (LUT) instruction that converts compressed 4-bit tensor values into FP4, FP6, or FP8 formats immediately before matrix multiplication — reducing preprocessing overhead and improving Tensor Core utilization. CDNA 5 also adds support for the E5M3 floating-point format, using the unused sign bit to extend range with an extra exponent bit for better accuracy at scale.
These aren't incremental improvements. On DeepSeek-V4-Flash, one of the leading open-weight models, AMD claims MI455X delivers up to 34× higher token throughput and 18× lower token cost compared to the previous-generation MI355X.
But the real story isn't just the GPU. It's the rack.

The Rack Is the Computer
Earlier AI systems were built by connecting servers containing eight GPUs each. AMD Helios shifts the boundary. Built on Open Compute Project (OCP) Open Rack Wide standards — developed in partnership with Meta — Helios treats 72 GPUs as a unified system. Each rack integrates 18 EPYC "Venice" processors (up to 256 cores each), up to 36 TB of DDR5 memory, and direct liquid cooling, with tray-level serviceability designed for hyperscale operations.
The networking layer is where AMD's open-standards bet becomes concrete. Within the rack, Ultra Accelerator Link over Ethernet (UALoE) delivers 260 TB/s of scale-up bandwidth. Across racks, the Ultra Ethernet Consortium (UEC) framework and Open ESUN standard provide scale-out networking without proprietary fabrics. This is a deliberate contrast to NVIDIA's NVLink/NVSwitch approach, which ties customers to a proprietary interconnect stack.
The software story has also matured. ROCm now natively supports PyTorch, TensorFlow, and JAX. At the rack level, AMD Fabric Manager handles UALoE routing and virtual pod configuration, AMD Rack Infrastructure Manager provides Redfish-based telemetry, and native Slurm and Kubernetes integration means Helios racks can be commissioned as part of larger AI factories rather than standalone boxes.

The Competitive Equation
AMD is positioning Helios directly against NVIDIA's Vera Rubin NVL72, which entered full production in May and began shipping to cloud providers in July. On paper, AMD claims Helios delivers 15% more AI compute, 50% more HBM capacity, and 50% more scale-out bandwidth than Vera Rubin.
The context matters. NVIDIA's Vera Rubin ramp is reshaping hyperscaler spending — CoreWeave reported its NVL72 racks deliver ten times the token output of Blackwell. Fireworks AI raised $1.5 billion at a $17.5 billion valuation, riding the same inference demand wave that fills Rubin-equipped racks. AMD is entering a market where the demand signal is unambiguous but the supply side remains concentrated.
AMD's timing is strategic. Anthropic confirmed on August 5 that it is building an in-house silicon team, OpenAI unveiled its custom Jalapeño inference chip co-designed with Broadcom in June, and every major lab is exploring alternatives to NVIDIA dependency. AMD isn't just competing with one company — it's positioning Helios as the open alternative in a market where every hyperscaler wants a second source.
The open-standards architecture is the differentiator. OCP ORW, UEC, and Open ESUN mean customers aren't locked into AMD's fabric. They can scale with Ethernet-based networking from multiple vendors. For cloud providers who've watched NVIDIA's margins and want negotiating leverage, that matters.
Builder Impact
For AI builders and infrastructure teams, Helios changes several equations simultaneously.
First, the HBM4 capacity jump — 31 TB per rack — means larger models can stay resident in GPU memory without sharding across racks. For teams deploying models in the 100B-400B parameter range, this reduces the engineering complexity of distributed inference and training.
Second, the open networking stack means infrastructure teams can use familiar Ethernet tooling and supply chains rather than learning a proprietary fabric. AMD Pensando AI-NICs and UALoE run over standard Ethernet — the same switches, cables, and monitoring tools that data center teams already operate.
Third, ROCm's maturation to "day one" support for PyTorch, JAX, and TensorFlow removes the historical barrier that kept many teams on CUDA. If the software stack is genuinely competitive — and AMD's DeepSeek-V4-Flash benchmarks suggest it might be — the CUDA moat starts to look shallower.
The caveat is real-world performance. AMD's 34× throughput claim for DeepSeek-V4-Flash is a vendor benchmark. Production workloads with different model architectures, batch sizes, and serving patterns will tell a more nuanced story. The key question builders should ask: can Helios deliver competitive price-performance on their workloads, not just on AMD's chosen benchmarks?
What's Next
AMD confirmed that Helios is being adopted across "AI leaders, cloud partners, and infrastructure partners," with AWS, Google Cloud, Microsoft, and OCI all named as early deployers alongside NVIDIA Cloud Partners CoreWeave, Lambda, Nebius, and Nscale. Dell, HPE, Lenovo, and Supermicro are building Helios-based systems.
The broader market signal is clear: the AI hardware market is fragmenting from a single-supplier dynamic into a multi-vendor landscape. NVIDIA still dominates — Vera Rubin NVL72 is shipping in volume, and the Blackwell installed base is enormous. But Anthropic is building custom silicon, OpenAI has Jalapeño, Google has TPUs, Amazon has Trainium, and now AMD has a credible rack-scale platform with open standards.
For the AI industry, this is healthy. Competition at the hardware layer drives down inference costs, which expands the market for AI applications. When compute is cheaper, more experiments run, more products ship, and more builders can afford to deploy at scale.
AMD's execution track record will determine whether Helios is a turning point or a footnote. The company has shipped competitive server CPUs for years with EPYC. If it can replicate that in GPU-accelerated AI with the same cadence — annual silicon updates, reliable software, enterprise support — the AI infrastructure market looks fundamentally different in 2027 than it did in 2025.
The rack is no longer where the AI system lives. The rack is the AI system. AMD just made that vision open.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
6 August 2026
6 August 2026
Sources
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.




