B200 vs H100 is a comparison of two Nvidia architectures: the B200 is the Blackwell-generation successor to the Hopper-based H100. Blackwell puts two dies and 208 billion transistors into one GPU (the H100 has 80 billion), adds FP4 math, and doubles NVLink bandwidth per GPU from 900 GB/s to 1.8 TB/s. Nvidia publishes most B200 numbers for systems rather than for one chip, so a fair comparison has to divide those system figures, as done below.

On this page

The short version

H100 (Hopper) B200 (Blackwell)
Announced March 22, 2022 March 18, 2024
Transistors 80 billion 208 billion
Process TSMC 4N TSMC 4NP
Dies per GPU 1 2, linked at 10 TB/s
Lowest Tensor Core precision FP8 FP4
NVLink generation, per GPU 4th gen, 900 GB/s 5th gen, 1.8 TB/s
Memory type HBM3 (SXM) HBM3e

Sources: Nvidia’s Hopper and Blackwell launch releases, the Hopper architecture post, the Blackwell architecture page and the HGX page.

What Blackwell changes in the chip

Two dies acting as one GPU. Nvidia describes Blackwell as “two reticle-limited dies connected by a 10 terabytes per second (TB/s) chip-to-chip interconnect in a unified single GPU.” Reticle-limited means each die is as large as the lithography mask allows. The H100, by comparison, is a single 814 mm² die.

FP4 through a second-generation Transformer Engine. Hopper’s Transformer Engine switches between FP8 and 16-bit formats. Blackwell’s version adds what Nvidia calls “micro-tensor scaling” to enable 4-bit floating point (FP4) AI. Halving the bits again lets the same memory hold larger models, and Nvidia’s HGX B200 figures show FP4 at twice the FP8 rate, for workloads that tolerate the lower precision.

Fifth-generation NVLink. Each B200 gets 1.8 TB/s of NVLink bandwidth, twice the H100 SXM’s 900 GB/s. Nvidia says fifth-generation NVLink can scale to 576 GPUs and that a 72-GPU domain (the GB200 NVL72 rack) has 130 TB/s of GPU bandwidth.

New engines. Blackwell adds a Decompression Engine for formats such as LZ4, Snappy and Deflate, aimed at database work, and a Reliability, Availability and Serviceability (RAS) engine for predictive fault detection. Nvidia calls Blackwell “the first TEE-I/O capable GPU in the industry,” extending the confidential-computing feature Hopper introduced.

Specs: how Nvidia documents each one

Nvidia lists the H100 as a single GPU. For the B200 it lists the eight-GPU HGX B200 board, which Nvidia’s HGX page says is “shipping now,” and the GB200 Grace Blackwell Superchip, which pairs two B200 GPUs with one Grace CPU. Tensor Core figures are with sparsity, as Nvidia states them; for HGX B200 Nvidia notes that dense is half the sparse figure, and gives FP4 as sparse | dense.

Spec (Nvidia figures) H100 SXM (1 GPU) HGX B200 (8 GPUs) GB200 Superchip (2 GPUs + Grace)
FP4 Tensor Core — 144 | 72 PFLOPS 40 | 20 PFLOPS
FP8 Tensor Core (sparse) 3,958 TFLOPS 72 PFLOPS 20 PFLOPS
FP16/BF16 Tensor Core (sparse) 1,979 TFLOPS 36 PFLOPS 10 PFLOPS
TF32 Tensor Core (sparse) 989 TFLOPS 18 PFLOPS 5 PFLOPS
FP64 / FP64 Tensor Core 34 / 67 TFLOPS 296 TFLOPS 80 TFLOPS
GPU memory 80 GB HBM3 1.4 TB 372 GB HBM3E
Memory bandwidth 3.35 TB/s 64 TB/s (DGX B200) 16 TB/s
NVLink per GPU 900 GB/s 1.8 TB/s 1.8 TB/s
Max power Up to 700 W ~14.3 kW (DGX B200 system) —

Per GPU, the arithmetic

Dividing Nvidia’s system figures by the number of GPUs gives a per-B200 view. These are our calculations from Nvidia’s numbers, not figures Nvidia lists per chip:

  • FP8 (sparse): 72 ÷ 8 = 9 PFLOPS per B200 in HGX B200, and 20 ÷ 2 = 10 PFLOPS per B200 in GB200, against 3.958 PFLOPS for the H100 SXM. That is about 2.3x and 2.5x.
  • Memory: the DGX B200 lists 1,440 GB across eight GPUs, or 180 GB each; the GB200 Superchip lists 372 GB across two, or 186 GB each. The H100 SXM has 80 GB.
  • Memory bandwidth: 64 ÷ 8 and 16 ÷ 2 both come to 8 TB/s per GPU, about 2.4x the H100 SXM’s 3.35 TB/s.
  • FP64: 296 ÷ 8 = 37 TFLOPS in HGX B200 and 80 ÷ 2 = 40 TFLOPS in GB200, lower than the H100 SXM’s 67 TFLOPS FP64 Tensor Core figure. Blackwell’s gains are in AI precisions, not double precision.

The gap between the HGX B200 and GB200 per-GPU figures matters when someone quotes “B200 specs”: the same B200 is rated differently depending on the system it sits in. Nvidia’s HGX page also gives 1.4 TB as the HGX B200 memory total, while the DGX B200 page lists 1,440 GB; each figure is used here as Nvidia states it for that product.

What Nvidia claims at the system level

Nvidia’s comparisons are system against system, and they are Nvidia’s claims:

  • The DGX B200 delivers “3X the training performance and 15X the inference performance of previous-generation systems.”
  • The GB200 NVL72 rack provides “up to a 30x performance increase compared to the same number of NVIDIA H100 Tensor Core GPUs for LLM inference workloads, and reduces cost and energy consumption by up to 25x,” according to the Blackwell launch release.

Those multiples are measured on whole systems running large models. The per-GPU FP8 ratio calculated above is about 2.3–2.5x, so most of the claimed inference gain comes from the system: FP4, more memory per GPU and the larger NVLink domain.

Where B300 fits

The HGX B300 board uses eight Blackwell Ultra GPUs. Nvidia lists it at the same 72 PFLOPS of sparse FP8 as HGX B200, a higher dense FP4 figure (108 vs 72 PFLOPS), 2.1 TB of memory against 1.4 TB, and twice the networking bandwidth (1.6 TB/s vs 0.8 TB/s). Its FP64 figure drops to 10 TFLOPS and INT8 to 3 POPS for the eight-GPU board, so it gives up HPC and integer throughput in favor of AI formats.

For a running list of accelerator specs across Nvidia and AMD, see the AI chips catalog; per-hour cloud rates for H100 and B200 instances are on the GPU prices page.

Sources