Microsoft Maia 200
Microsoft Maia 200 is Microsoft’s in-house inference accelerator, introduced on January 26, 2026.
- Vendor
- Microsoft
- Architecture
- Maia 200, TSMC 3nm
- Memory
- 216 GB HBM3e
- Memory bandwidth (TB/s)
- 7 TB/s
- FP8 (TFLOPS)
- 5,000 TFLOPS
- FP4 (PFLOPS)
- 10 PFLOPS
- Power (W)
- 750 W
- Form factor
- Azure servers (four accelerators per tray)
- Status
- Shipping
Maia 200 has over 140 billion transistors on TSMC 3nm, 216 GB of HBM3e at 7 TB/s and 272 MB of on-chip SRAM within a 750 W SoC TDP. Microsoft says it is deployed in its US Central Azure region near Des Moines, Iowa, with US West 3 near Phoenix next, and serves models including OpenAI’s GPT-5.2. Clusters scale to 6,144 accelerators over Ethernet.
- Memory (GB)
- 216 GB
- Precision note (sparse or dense)
- Microsoft states "over 10 petaFLOPS" FP4 and "over 5 petaFLOPS" FP8 per chip, with no sparse/dense label.
- Launch
- Introduced January 26, 2026
- Vendor source
- blogs.microsoft.com
- Checked
Questions
How much memory does Microsoft Maia 200 have?
Maia 200 has 216 GB of HBM3e at 7 TB/s and 272 MB of on-chip SRAM.
What is Maia 200’s performance?
Microsoft states over 10 petaFLOPS FP4 and over 5 petaFLOPS FP8 per chip, with no sparse or dense label.
Where is Maia 200 deployed?
Microsoft says it is deployed in its US Central Azure region near Des Moines, Iowa, with US West 3 near Phoenix next.
What models run on Maia 200?
Microsoft says Maia 200 serves models including OpenAI’s GPT-5.2.
Source
Every figure is taken from the vendor’s own datasheet, product page, technical documentation or announcement, linked on each chip with the date it was last checked. How we pick and check sources: editorial standards.