The GB200 NVL72 is a single liquid-cooled Nvidia rack that holds 36 Grace CPUs and 72 Blackwell GPUs and links all 72 GPUs in one NVLink domain, so that, in Nvidia’s words, the rack “acts as a single, massive GPU.” Nvidia lists the GB200 NVL72 at 13.4 TB of HBM3E, 130 TB/s of NVLink bandwidth and 720 petaFLOPS of FP8 compute with sparsity. Nvidia’s DGX GB200 user guide puts the rack’s power consumption at about 120 kW.
On this page
What is inside the rack
The basic unit is the GB200 Grace Blackwell Superchip: one Grace CPU and two B200 GPUs, connected by NVLink-C2C at 900 GB/s, which Nvidia says gives “coherent access to a unified memory space.” The rack is built from trays of these superchips plus trays of NVLink switches. Nvidia’s DGX GB200 user guide lists:
- 18 compute trays (1U each), each with 2 Grace CPUs and 4 Blackwell GPUs, so two superchips per tray
- 9 NVLink switch trays (1U each), each with 2 NVLink switch chips
- 2 top-of-rack switches for management
- 8 power shelves with N+N redundancy, each using six 5.5 kW power supplies
Per compute tray, the guide lists four ConnectX-7 400G network cards and two BlueField-3 DPUs for traffic leaving the rack. Nvidia’s technical blog adds that each compute tray delivers 80 petaFLOPS of AI performance and 1.7 TB of fast memory, and that the switch trays are joined to the compute trays by a copper cable cartridge, which Nvidia chose “for operational simplicity.”
The NVLink domain
NVLink is Nvidia’s GPU-to-GPU link, and the size of an NVLink domain sets how many GPUs can exchange data over it directly rather than over the data-center network. On an eight-GPU HGX server the domain is eight GPUs. In the GB200 NVL72 it is 72.
Each Blackwell GPU has 1.8 TB/s of bidirectional NVLink bandwidth. The nine switch trays connect all 72 GPUs to each other for a total of 130 TB/s. Nvidia’s technical blog says the NVLink Switch System can extend a domain to “up to 576 GPUs” with “over 1 PB/s total bandwidth.”
Why the domain size matters: a model too large for one GPU is split across many, and the pieces exchange data for every token they process. Inside one NVLink domain that exchange runs at 1.8 TB/s per GPU; between servers it runs over network cards, which in this rack are 400 Gb/s each.
GB200 NVL72 totals
Nvidia’s product page lists the following. Tensor Core figures are with sparsity; FP4 is shown as sparse | dense, as Nvidia gives it.
| Spec | GB200 NVL72 | GB200 Superchip |
|---|---|---|
| Configuration | 36 Grace CPUs, 72 Blackwell GPUs | 1 Grace CPU, 2 Blackwell GPUs |
| NVFP4 Tensor Core | 1,440 | 720 PFLOPS | 40 | 20 PFLOPS |
| FP8/FP6 Tensor Core | 720 PFLOPS | 20 PFLOPS |
| FP16/BF16 Tensor Core | 360 PFLOPS | 10 PFLOPS |
| FP64 / FP64 Tensor Core | 2,880 TFLOPS | 80 TFLOPS |
| GPU memory | 13.4 TB HBM3E | 372 GB HBM3E |
| GPU memory bandwidth | 576 TB/s | 16 TB/s |
| NVLink bandwidth | 130 TB/s | 3.6 TB/s |
| CPU cores | 2,592 Arm Neoverse V2 | 72 Arm Neoverse V2 |
| CPU memory | 17 TB LPDDR5X, 14 TB/s | Up to 480 GB LPDDR5X, up to 512 GB/s |
Nvidia’s claims for the rack, against the same number of H100 GPUs: “30x faster real-time trillion-parameter large language model (LLM) inference,” “4x faster training for large language models at scale,” and “25x more performance at the same power.”
Power and cooling
The GB200 NVL72 is liquid-cooled. Nvidia’s DGX GB200 user guide gives rack power consumption of about 120 kW, delivered by eight power shelves. For comparison, Nvidia lists the eight-GPU DGX B200 system at about 14.3 kW. Facilities that host these racks have to supply that power and liquid cooling per rack; projects being built for them are tracked on our AI data centers page.
GB300 NVL72: the successor
The GB300 NVL72 keeps the rack shape, 72 GPUs and 36 Grace CPUs, and swaps in Blackwell Ultra GPUs. Nvidia describes it as a “fully liquid-cooled, rack-scale architecture.” Compared on Nvidia’s own product pages (Tensor Core figures with sparsity):
| Spec | GB200 NVL72 | GB300 NVL72 |
|---|---|---|
| GPUs | 72 Blackwell | 72 Blackwell Ultra |
| FP4 Tensor Core (sparse) | 1,440 PFLOPS | 1,440 PFLOPS |
| FP8/FP6 Tensor Core | 720 PFLOPS | 720 PFLOPS |
| INT8 Tensor Core | 720 POPS | 24 POPS |
| FP64 / FP64 Tensor Core | 2,880 TFLOPS | 100 TFLOPS |
| GPU memory | 13.4 TB HBM3E | 20 TB |
| GPU memory bandwidth | 576 TB/s | Up to 576 TB/s |
| NVLink bandwidth | 130 TB/s | 130 TB/s |
| Network per GPU | ConnectX-7, 400 Gb/s | ConnectX-8, 800 Gb/s |
The main change is memory: about 50% more GPU memory in the same rack (Nvidia describes “1.5x larger HBM3E memory”), plus twice the network bandwidth per GPU and what Nvidia calls “2X the attention-layer acceleration” in Blackwell Ultra Tensor Cores, while FP64 and INT8 throughput fall sharply. Nvidia claims the GB300 NVL72 gives a “10x boost in user responsiveness” and a “5x improvement in throughput” per megawatt against Hopper, which it combines into “a 50x leap in overall AI factory output.”
After GB300, Nvidia’s next rack in the same 72-GPU format is Vera Rubin NVL72. Specs for all three are kept side by side in the AI chips catalog, and cloud rental rates for GB200 capacity, where providers publish them, are on GPU prices.




