The AMD Instinct MI355X is AMD’s top CDNA 4 data-center accelerator, launched on June 12, 2025, with 288 GB of HBM3E memory at 8 TB/s and a typical board power of 1,400 W. AMD rates the AMD MI355X at 5 petaFLOPS of dense FP8 (10.1 with structured sparsity) and 10.1 petaFLOPS of MXFP4 and MXFP6. As of October 2026 it sits below AMD’s MI400 series, which AMD launched on July 23, 2026.
On this page
What the MI355X is
The MI355X belongs to AMD’s Instinct MI350 series, built on the CDNA 4 architecture. AMD lists it on TSMC 3nm and 6nm FinFET processes, with 256 compute units, 16,384 stream processors and 1,024 matrix cores running at a peak engine clock of 2,400 MHz. It ships as an OAM module (the Open Compute Project accelerator module form factor), and AMD lists both passive and active cooling options.
The series has two members. The MI350X has the same 256 compute units but runs at a 2,200 MHz peak clock with a 1,000 W typical board power and passive OAM cooling. The MI355X raises the clock to 2,400 MHz and the board power to 1,400 W; its FP8 figures are about 9–10% higher (5 vs 4.6 petaFLOPS dense, 10.1 vs 9.2 with sparsity). Memory is the same on both: 288 GB of HBM3E at 8 TB/s through an 8,192-bit interface, according to AMD’s MI350X page, which also gives the chip as 185 billion transistors.
New precisions in CDNA 4
CDNA 4 adds the OCP microscaling formats MXFP6 and MXFP4, which AMD’s MI300X and MI325X pages do not list. AMD rates the MI355X at 10.1 petaFLOPS for both, the same rate as its FP8 figure with sparsity. Because FP6 runs at the FP4 rate, a model quantized to six bits gets the same peak throughput as one quantized to four bits while keeping more precision.
AMD also lists FP16 and BF16 matrix throughput at 2.5 petaFLOPS dense and 5 petaFLOPS with structured sparsity, about twice the 1.3 petaFLOPS dense FP16 figure AMD gives for MI300X and MI325X.
MI355X vs MI300X and MI325X
All figures below are from AMD’s product pages as of October 2026. AMD gives FP8 both dense and with structured sparsity; both are shown.
| Spec | MI300X | MI325X | MI355X |
|---|---|---|---|
| Architecture | CDNA 3 | CDNA 3 | CDNA 4 |
| Process | TSMC 5nm | 6nm | TSMC 5nm | 6nm | TSMC 3nm | 6nm |
| Compute units | 304 | 304 | 256 |
| Memory | 192 GB HBM3 | 256 GB HBM3E | 288 GB HBM3E |
| Memory bandwidth | 5.3 TB/s | 6 TB/s | 8 TB/s |
| FP16 (dense) | 1.3 PFLOPS | 1.3 PFLOPS | 2.5 PFLOPS |
| FP8 (dense / sparse) | 2.61 / 5.22 PFLOPS | 2.61 / 5.22 PFLOPS | 5 / 10.1 PFLOPS |
| MXFP4 / MXFP6 | — | — | 10.1 PFLOPS |
| FP64 | 81.7 TFLOPS | 81.7 TFLOPS | 78.6 TFLOPS |
| Typical board power | 750 W peak | 1,000 W peak | 1,400 W |
| Form factor | OAM | OAM | OAM |
Three things stand out. Memory grew in every step, from 192 GB to 256 GB to 288 GB, and bandwidth from 5.3 to 8 TB/s. FP8 roughly doubled from MI325X to MI355X while the compute-unit count fell from 304 to 256. And FP64 stayed roughly flat, so the gain is in AI precisions, not double-precision math. AMD still claims “up to 2.1X the HPC performance vs. competitive accelerators” for the MI355X.
How AMD positions it against Nvidia
AMD’s product page claims the MI355X offers “up to 2.2X the AI performance vs. competitive accelerators” and “1.6X Memory Capacity vs. competitive accelerators.” These are AMD’s claims, and the page does not name the comparison product.
The memory figure can be set against Nvidia’s own numbers. Nvidia lists the eight-GPU DGX B200 with 1,440 GB of HBM3e, or 180 GB per GPU; 288 GB is 1.6 times that. On FP8 with sparsity, AMD’s 10.1 petaFLOPS compares with 9 petaFLOPS per GPU from Nvidia’s HGX B200 board figure (72 petaFLOPS across eight GPUs, our division) and 10 petaFLOPS per GPU in the GB200 Superchip (20 petaFLOPS across two). Peak figures are each vendor’s own, so the AI chips catalog keeps each vendor’s sparse and dense labels as published.
8-GPU platform and interconnect
AMD lists seven Infinity Fabric links per MI355X at 153 GB/s each for GPU-to-GPU traffic, a PCIe 5.0 x16 host interface, and 128 GB/s of scale-out bandwidth per GPU in an eight-GPU platform. Eight MI355X accelerators hold 2.3 TB of HBM3E in one server (our multiplication of AMD’s per-GPU figure).
Where the MI400 series leaves the MI355X
On July 23, 2026, AMD launched the Instinct MI400 series with two products: the MI455X, for frontier-model training and high-volume inference, and the MI430X, for sovereign AI and HPC, which AMD rates at “up to 288 TFLOPS of hardware-based FP64 performance.” AMD says the series uses HBM4 memory, and the MI455X powers AMD’s Helios rack-scale system. The MI355X is AMD’s HBM3E-based accelerator one step below the MI400 line.
Cloud rental rates for MI355X and other AMD Instinct GPUs, where providers publish them, are on the GPU prices page.
Sources
- AMD, Instinct MI355X product page
- AMD, Instinct MI350X product page
- AMD, Instinct MI325X product page
- AMD, Instinct MI300X product page
- AMD, AMD Instinct MI400 Series update, July 23, 2026
- Nvidia, DGX B200 product page
- Nvidia, HGX platform specifications
- Nvidia, GB200 NVL72 product page




