The Nvidia H100 is Nvidia’s Hopper-generation data-center GPU, announced on March 22, 2022, and built to train and run large AI models. The SXM version of the Nvidia H100 carries 80 GB of HBM3 memory at 3.35 TB/s and draws up to 700 W; the PCIe-based H100 NVL carries 94 GB at 3.9 TB/s. Nvidia rates the SXM part at 3,958 teraFLOPS of FP8 Tensor Core throughput, a figure it lists with sparsity.

On this page

What the H100 is

The H100 is the first GPU on Nvidia’s Hopper architecture, named after the computer scientist Grace Hopper. According to Nvidia’s launch release, it has 80 billion transistors and is made on a TSMC 4N process customized for Nvidia. Nvidia said at launch that the H100 would be available starting in the third quarter of 2022.

Nvidia’s architecture write-up gives the layout: the full GH100 chip has 144 streaming multiprocessors (SMs), of which the H100 SXM5 enables 132 and the H100 PCIe enables 114. Both share a 50 MB L2 cache, and the die measures 814 mm².

Two features define Hopper for AI work:

  • Transformer Engine. Nvidia describes it as hardware and software that “intelligently manages and dynamically chooses between FP8 and 16-bit calculations.” FP8 halves the bits per value compared with FP16, so more math fits through the same silicon and memory.
  • Fourth-generation NVLink. The SXM version connects to other GPUs at 900 GB/s, and Nvidia’s NVLink Switch System can join up to 256 GPUs.

H100 versions: SXM, PCIe and NVL

“H100” covers more than one product, and the numbers differ between them.

  • H100 SXM is a module that mounts on Nvidia’s HGX H100 boards with four or eight GPUs. It has the highest power budget (up to 700 W, configurable) and the full 900 GB/s NVLink.
  • H100 PCIe was the original card version. Nvidia’s architecture post lists it with HBM2e memory at more than 2 TB/s and a 350 W TDP.
  • H100 NVL is a dual-slot, air-cooled PCIe card with 94 GB of HBM3 at 3,938 GB/s and a 350–400 W configurable TDP. Pairs of cards can be bridged with NVLink at 600 GB/s. Nvidia’s product brief calls it “the most optimized platform for LLM Inferences,” and the card ships with a five-year Nvidia AI Enterprise subscription.

H100 specs at a glance

Figures below are from Nvidia’s H100 product page as of October 2026. Tensor Core rows are the numbers Nvidia lists with sparsity; dense throughput is lower.

Spec H100 SXM H100 NVL
FP64 34 TFLOPS 30 TFLOPS
FP64 Tensor Core 67 TFLOPS 60 TFLOPS
TF32 Tensor Core (sparse) 989 TFLOPS 835 TFLOPS
BF16 / FP16 Tensor Core (sparse) 1,979 TFLOPS 1,671 TFLOPS
FP8 Tensor Core (sparse) 3,958 TFLOPS 3,341 TFLOPS
INT8 Tensor Core (sparse) 3,958 TOPS 3,341 TOPS
GPU memory 80 GB 94 GB
Memory bandwidth 3.35 TB/s 3.9 TB/s
Max TDP Up to 700 W 350–400 W
Multi-Instance GPU Up to 7 × 10 GB Up to 7 × 12 GB
Interconnect NVLink 900 GB/s, PCIe Gen5 128 GB/s NVLink 600 GB/s, PCIe Gen5 128 GB/s
Server options HGX H100 (4 or 8 GPUs) Partner systems (1–8 GPUs)

Two details matter when reading these numbers. First, the sparsity figures assume a model whose weights have been pruned into a pattern the Tensor Cores can skip; a dense FP8 number from another vendor is not comparable with 3,958 teraFLOPS. Second, the H100 NVL has more memory and more bandwidth than the SXM part but lower compute, because it runs at roughly half the power.

What the H100 is used for

The H100 was designed for three kinds of work, and Nvidia’s own claims are organized the same way:

  • Training large language models. Nvidia claims “up to 4X faster training over the prior generation for GPT-3 (175B) models,” the prior generation being the A100.
  • Inference. Nvidia claims the H100 can “accelerate inference by up to 30X” on the largest models compared with A100, crediting the Transformer Engine and FP8.
  • High-performance computing. Nvidia states that the H100 “triples the floating-point operations per second (FLOPS) of double-precision Tensor Cores” and claims up to 7X higher performance for HPC applications. Its DPX instructions, which speed up dynamic-programming algorithms used in genomics and route planning, are rated by Nvidia at 7X over A100.

Beyond raw speed, two features shape how clouds sell the H100. Multi-Instance GPU (MIG) splits one card into up to seven isolated instances, which lets a provider rent a fraction of a GPU. Confidential computing, which Nvidia calls “a built-in security feature of the NVIDIA Hopper architecture,” keeps data and models encrypted while in use. Current hourly rental rates from cloud providers are tracked on our GPU prices page.

Where the H100 sits next to H200 and B200

H200 uses the same Hopper architecture and the same compute ratings: Nvidia lists the H200 SXM at 3,958 teraFLOPS FP8 with sparsity, identical to the H100 SXM. The change is memory. Nvidia calls the H200 “the first GPU to offer 141 gigabytes (GB) of HBM3e memory at 4.8 terabytes per second,” which it describes as “nearly double the capacity” of the H100 with 1.4X more memory bandwidth. For inference on large models, where the GPU spends much of its time waiting on memory, that difference shows up directly in throughput.

B200 is the next architecture, Blackwell. Nvidia’s Blackwell GPU has 208 billion transistors across two dies joined by a 10 TB/s chip-to-chip link, and adds FP4 support through a second-generation Transformer Engine. Nvidia publishes B200 figures mainly at the system level: an HGX B200 board with eight GPUs is rated at 72 petaFLOPS of FP8 with sparsity, and the DGX B200 system carries 1,440 GB of HBM3e in total. Nvidia claims the DGX B200 delivers “3X the training performance and 15X the inference performance” of previous-generation systems.

Side-by-side figures for every Nvidia and AMD data-center accelerator are in our AI chips catalog.

Sources