Nvidia Rubin is the data-center GPU generation that follows Blackwell, sold as the Vera Rubin platform: Rubin GPUs paired with Nvidia’s Vera CPU. Nvidia lists each Rubin GPU with 288 GB of HBM4 at 19.2 TB/s and 50 petaFLOPS of NVFP4 inference compute. In its second-quarter fiscal 2027 results on August 26, 2026, Nvidia said the platform is “ramping into full production with racks running at partners.” Everything below is limited to what Nvidia has announced, as of October 2026.
On this page
Timeline of announcements
| Date | Where | What Nvidia announced |
|---|---|---|
| January 5, 2026 | CES | Rubin platform of six chips; Vera Rubin NVL72 and HGX Rubin NVL8; “in full production,” with partner availability in the second half of 2026 |
| March 16, 2026 | GTC | Vera Rubin platform expanded to seven chips with the Groq 3 LPU; Vera CPU launched as a standalone product |
| August 24, 2026 | Hot Chips | Groq 3 LPX “now in full production” |
| August 26, 2026 | Q2 FY2027 results | Vera Rubin “ramping into full production with racks running at partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius” |
The platform is named after Vera Florence Cooper Rubin, the American astronomer, according to Nvidia’s CES release.
The chips in the platform
At CES Nvidia listed six chips designed together: the Vera CPU, the Rubin GPU, the NVLink 6 Switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU and the Spectrum-6 Ethernet switch. At GTC in March it added a seventh, the Groq 3 LPU, which comes in a separate LPX rack.
Rubin GPU specs
Nvidia’s Vera Rubin NVL72 product page gives the figures below. Nvidia marks the NVFP4 inference figures as sparse and the training, FP8 and lower-precision rows as dense.
| Spec | Rubin GPU | Vera Rubin Superchip | Vera Rubin NVL72 |
|---|---|---|---|
| Configuration | 1 Rubin GPU | 2 Rubin GPUs + 1 Vera CPU | 72 Rubin GPUs + 36 Vera CPUs |
| NVFP4 inference (sparse) | 50 PFLOPS | 100 PFLOPS | 3,600 PFLOPS |
| NVFP4 training (dense) | 35 PFLOPS | 70 PFLOPS | 2,520 PFLOPS |
| FP8/FP6 training (dense) | 17.5 PFLOPS | 35 PFLOPS | 1,260 PFLOPS |
| GPU memory | 288 GB HBM4 | 576 GB HBM4 | 20.7 TB HBM4 |
| Memory bandwidth | 19.2 TB/s | 38.5 TB/s | 1,400 TB/s |
| NVLink 6 bandwidth | 3 TB/s | 6 TB/s | 216 TB/s |
| CPU | — | 88 Olympus cores, up to 1.5 TB LPDDR5X | 3,168 Olympus cores, up to 54 TB LPDDR5X |
Nvidia’s January release described NVLink 6 at 3.6 TB/s per GPU and 260 TB/s per NVL72 rack; the product page, as of October 2026, lists 3 TB/s and 216 TB/s. The table uses the product page.
Nvidia also describes a third-generation Transformer Engine “with adaptive compression” in the Rubin GPU. HBM4 is the memory generation that doubles the interface width to 2,048 bits per stack under the JEDEC standard, which is where most of the bandwidth gain over HBM3e comes from.
The Vera CPU
Vera is Nvidia’s Arm-compatible CPU that follows Grace. Nvidia’s March 16, 2026 release lists 88 custom Olympus cores, each running two tasks through what Nvidia calls Spatial Multithreading, LPDDR5X memory with up to 1.2 TB/s of bandwidth, and 1.8 TB/s of NVLink-C2C bandwidth to the GPUs, which Nvidia puts at 7x PCIe Gen 6.
Vera is also sold on its own. Nvidia describes a Vera CPU rack with 256 liquid-cooled CPUs and says standalone Vera will be available from partners in the second half of 2026. Named adopters in that release include Alibaba Cloud, ByteDance, Meta and Oracle Cloud Infrastructure, along with CoreWeave, Lambda, Nebius and Nscale. In its Q2 FY2027 results Nvidia called Vera “the first CPU built for AI agents.”
Vera Rubin NVL72 and HGX Rubin NVL8
The Vera Rubin NVL72 is the rack-scale system: 72 Rubin GPUs and 36 Vera CPUs connected by NVLink 6, following the layout of the Blackwell-generation GB200 NVL72. Nvidia claims, compared with Blackwell, “one-tenth the cost per million tokens” for inference, “up to 10x more tokens per megawatt,” and training with “one-fourth the number of GPUs” for mixture-of-experts models.
For eight-GPU servers, Nvidia lists the HGX Vera Rubin NVL8 (eight Rubin SXM GPUs with a single Vera CPU) at 400 PFLOPS of NVFP4 inference, 2 TB of HBM4 at 130 TB/s, and sixth-generation NVLink with 24 TB/s of switch bandwidth. An HGX Rubin NVL8 variant pairs the same eight GPUs with an x86 CPU.
Groq 3 LPX: the inference extension
The Groq 3 LPX is a rack of Groq 3 LPU processors that Nvidia positions alongside Vera Rubin NVL72 for fast token generation. Nvidia’s GTC release says “the LPX rack with 256 LPU processors features 128GB of on-chip SRAM and 640 TB/s of scale-up bandwidth,” and claims the combination delivers “up to 35x higher inference throughput per megawatt.”
On August 24, 2026, Nvidia said Groq 3 LPX is in full production and “extends the inference performance of Vera Rubin NVL72 by dramatically increasing the rate of token generation.” Nvidia reported 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context in Artificial Analysis benchmarking, and named Nebius as “the first AI cloud to adopt NVIDIA Groq 3 LPX.”
What “full production” means here
Nvidia used “in full production” for Rubin at CES in January, while giving partner availability as the second half of 2026. The August results moved the wording to racks “running at partners,” with five clouds named. Where those racks are being installed is tracked on our AI data centers page, and Rubin’s specs sit next to Blackwell and AMD parts in the AI chips catalog.
Sources
- Nvidia, NVIDIA Kicks Off the Next Generation of AI With Rubin — Six New Chips, One Incredible AI Supercomputer, January 5, 2026
- Nvidia, NVIDIA Vera Rubin Opens Agentic AI Frontier, March 16, 2026
- Nvidia, NVIDIA Launches Vera CPU, Purpose-Built for Agentic AI, March 16, 2026
- Nvidia, NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI, August 24, 2026
- Nvidia, NVIDIA Announces Financial Results for Second Quarter Fiscal 2027, August 26, 2026
- Nvidia, Vera Rubin NVL72 product page
- Nvidia, HGX platform specifications
- JEDEC, JESD270-4 HBM4 standard release, April 16, 2025




