Microsoft Maia 200

Microsoft Maia 200 is Microsoft’s in-house inference accelerator, introduced on January 26, 2026.

Vendor
Microsoft
Architecture
Maia 200, TSMC 3nm
Memory
216 GB HBM3e
Memory bandwidth (TB/s)
7 TB/s
FP8 (TFLOPS)
5,000 TFLOPS
FP4 (PFLOPS)
10 PFLOPS
Power (W)
750 W
Form factor
Azure servers (four accelerators per tray)
Status
Shipping

Maia 200 has over 140 billion transistors on TSMC 3nm, 216 GB of HBM3e at 7 TB/s and 272 MB of on-chip SRAM within a 750 W SoC TDP. Microsoft says it is deployed in its US Central Azure region near Des Moines, Iowa, with US West 3 near Phoenix next, and serves models including OpenAI’s GPT-5.2. Clusters scale to 6,144 accelerators over Ethernet.

Memory (GB)
216 GB
Precision note (sparse or dense)
Microsoft states "over 10 petaFLOPS" FP4 and "over 5 petaFLOPS" FP8 per chip, with no sparse/dense label.
Launch
Introduced January 26, 2026
Vendor source
blogs.microsoft.com
Checked

Questions

How much memory does Microsoft Maia 200 have?
Maia 200 has 216 GB of HBM3e at 7 TB/s and 272 MB of on-chip SRAM.
What is Maia 200’s performance?
Microsoft states over 10 petaFLOPS FP4 and over 5 petaFLOPS FP8 per chip, with no sparse or dense label.
Where is Maia 200 deployed?
Microsoft says it is deployed in its US Central Azure region near Des Moines, Iowa, with US West 3 near Phoenix next.
What models run on Maia 200?
Microsoft says Maia 200 serves models including OpenAI’s GPT-5.2.

Source

Every figure is taken from the vendor’s own datasheet, product page, technical documentation or announcement, linked on each chip with the date it was last checked. How we pick and check sources: editorial standards.

The Model Press

What are you looking for?

Search by headline, topic or keyword.