Skip to content

Microsoft deploys AMD Helios rack-scale AI system on Azure for frontier-model inference

· by Pondero Newsdesk

The short version

Microsoft became the first named hyperscaler to publicly commit to AMD Helios at scale, agreeing to run the new rack-scale platform on Azure for frontier-model inference. Shipments begin H2 2026.

Microsoft deploys AMD Helios rack-scale AI system on Azure for frontier-model inference

Microsoft became the first named hyperscaler to publicly commit to AMD Helios at scale, agreeing on July 20 to run the new platform on Azure for frontier-model inference. The deal placed Azure at the center of AMD's commercial launch for Helios, which is AMD's first integrated rack-scale product and its most direct challenge yet to Nvidia's dominance of the cloud AI compute market.

What

Helios combines AMD Instinct MI455X GPUs, sixth-generation EPYC Venice CPUs, Pensando networking, and AMD ROCm software into a single integrated, liquid-cooled rack, per AMD's Newsroom. Each double-wide unit holds 72 MI455X accelerators with 31 TB of HBM4 memory and 1.4 exaFLOPS of FP8 throughput, per AIWeekly's coverage of the launch. AMD said shipments to Microsoft and other early customers begin in the second half of 2026.

Azure will surface the Helios hardware through an updated VM lineup. Two new EPYC Venice-only families arrive alongside the Helios deal: Azure HDv2, aimed at agentic AI workloads and data pipelines, and Azure HXv2, aimed at semiconductor design, per AMD. The Pensando DPU footprint on Azure also expands, with AMD and Microsoft integrating Pensando networking into Azure Boost to improve connection throughput at cloud scale.

AMD Chair and CEO Lisa Su called the Azure deployment "an important milestone as we deliver leadership compute solutions to Azure customers," per AMD's press release. Microsoft CEO Satya Nadella described customer demand for infrastructure "optimized for a wide range of workloads, from training and inference to data preparation, search, and reinforcement learning," per the same release.

Why it matters

Helios is AMD's direct answer to Nvidia's NVL72 form factor: a rack-scale unit that ships as one system rather than loose parts. Microsoft going on record as a named lighthouse customer is the public anchor AMD needed to move enterprise cloud buyers toward the platform. Meta, OpenAI, and Oracle are also reported early customers, per AIWeekly, but none made the same public commitment to running Helios for frontier-model inference on a named hyperscaler's public cloud.

For operators building AI workloads on Azure Foundry, the deployment has a concrete implication. AMD's press release names Azure Foundry Managed Compute as a distribution channel for Helios-backed capacity. Inference calls through Azure-hosted frontier models may run on MI455X silicon starting in H2 2026, without any change to the API surface those operators already use.

The HDv2 and HXv2 VM families carry their own signal. Naming semiconductor design as a distinct workload class for Azure HXv2 positions EPYC Venice against incumbents in EDA compute, a segment where AMD had not previously held a dedicated Azure VM. If AMD holds that position, it widens the CPU diversification story beyond AI inference.

What to watch next

Helios shipments begin H2 2026. General availability dates for the ND MI455X v7 VM, the HDv2 series, and the HXv2 series will show how quickly Azure makes the capacity accessible to production workloads. A second hyperscaler commitment from Google Cloud or AWS would confirm whether Helios has genuine multi-cloud traction or remains an Azure-anchored launch win.

Sources