AI chips compared: H100, B200, Rubin, MI355X, TPU and Trainium accelerator specs

Specs of data-center AI accelerators from every vendor, copied from each vendor’s own datasheet.

25 listings

NameVendorMemoryMemory bandwidth (TB/s)FP8 (TFLOPS)Status
AMD Instinct MI300XAMD192 GB HBM35.3 TB/s2,610 TFLOPSShipping
AMD Instinct MI325XAMD256 GB HBM3E6 TB/s2,610 TFLOPSShipping
AMD Instinct MI355XAMD288 GB HBM3E8 TB/s5,000 TFLOPSShipping
AMD Instinct MI430XAMD432 GB HBM423.3 TB/sAnnounced
AMD Instinct MI455XAMD432 GB HBM423.3 TB/s20,100 TFLOPSAnnounced
AWS Trainium2AWS96 GiB HBM2.9 TB/s1,299 TFLOPSShipping
AWS Trainium3AWS144 GB HBM3e4.9 TB/s2,517 TFLOPSShipping
Cerebras WSE-3T (CS-4)Cerebras44 GB on-wafer SRAM43,200 TB/sAnnounced
Google TPU 8iGoogle288 GB HBM8.601 TB/sAnnounced
Google TPU 8tGoogle216 GB HBM6.528 TB/sAnnounced
Google TPU7x (Ironwood)Google192 GB HBM7.37 TB/s4,614 TFLOPSShipping
Huawei Ascend 950DTHuawei144 GB HiZQ 2.0 (Huawei HBM)4 TB/s1,000 TFLOPSAnnounced
Huawei Ascend 960 (Atlas 960E SuperPoD)HuaweiUp to 1 PB HBM per Atlas 960E SuperPoD (4,096 NPUs)2,000 TFLOPSAnnounced
Intel Gaudi 3Intel128 GB HBM2e3.7 TB/s1,678 TFLOPSShipping
Meta MTIA 200Meta128 GB LPDDR5 (off-chip) + 256 MB SRAM0.205 TB/sShipping
Microsoft Maia 200Microsoft216 GB HBM3e7 TB/s5,000 TFLOPSShipping
Nvidia B200 (Blackwell)NvidiaUp to 192 GB HBM3E (180 GB per GPU in HGX B200)8 TB/s5,000 TFLOPSShipping
Nvidia B300 (Blackwell Ultra)Nvidia288 GB HBM3E8 TB/s5,000 TFLOPSShipping
Nvidia GB200 NVL72Nvidia13.4 TB HBM3E (72 GPUs)576 TB/s720,000 TFLOPSShipping
Nvidia GB300 NVL72Nvidia20 TB HBM3E (72 GPUs)576 TB/s720,000 TFLOPSShipping
Nvidia Groq 3 LPXNvidia128 GB SRAM per rack (256 LPUs)40,000 TB/sIn production
Nvidia H100 SXMNvidia80 GB HBM33.35 TB/s3,958 TFLOPSShipping
Nvidia H200 SXMNvidia141 GB HBM3e4.8 TB/s3,958 TFLOPSShipping
Nvidia Rubin GPUNvidia288 GB HBM422 TB/s17,500 TFLOPSIn production
Nvidia Vera Rubin NVL72Nvidia20.7 TB HBM4 (72 GPUs)1,580 TB/s1,260,000 TFLOPSIn production

This table lists AI accelerator specs for data-center chips from Nvidia, AMD, Google, AWS, Intel, Cerebras, Huawei, Microsoft and Meta: Hopper H100 and H200, Blackwell B200 and B300, GB200 and GB300 NVL72 racks, the Rubin GPU and Vera Rubin NVL72, Instinct MI300X through MI455X, TPU Ironwood and TPU 8, Trainium2 and Trainium3, Gaudi 3, Cerebras WSE-3T, Ascend, Maia 200 and MTIA.

Every number comes from the vendor’s own datasheet, product page, technical documentation or announcement, linked on each entry with the date it was checked. We copy figures as the vendor states them. Vendors quote FP8 and FP4 throughput either with sparsity or without it (dense); each entry’s precision note says which, because the two can differ by a factor of two. Where a vendor publishes no figure for a field, the field is left blank rather than estimated.

Some vendors publish specs per chip, others per board, rack or pod. The form factor column says which level a row describes, so rack-scale rows (NVL72, CS-4, Groq 3 LPX) are not directly comparable with single chips. Memory is in GB, bandwidth in TB/s, FP8 in TFLOPS and FP4 in PFLOPS.

The table is checked against vendor sources three times a week. For launch coverage see AI chip news, for cloud rental rates per GPU-hour see GPU prices, and for where these chips are being deployed see AI data centers. How we source data is described in our editorial standards.

Source

Every figure is taken from the vendor’s own datasheet, product page, technical documentation or announcement, linked on each chip with the date it was last checked. How we pick and check sources: editorial standards.

The Model Press

What are you looking for?

Search by headline, topic or keyword.