Nvidia said on August 24, 2026 that the Nvidia Groq 3 LPX is “now in full production,” in a release issued at Hot Chips. Nvidia describes the LPX as an inference accelerator that extends its Vera Rubin platform and is built to raise the rate of token generation for agentic AI systems.
On this page
Performance Nvidia claims
The figures below come from Nvidia’s release:
| Claim | Setting given by Nvidia |
|---|---|
| 3,400 output tokens per second | Gemma 4 31B with a 100,000-token context |
| 4x faster responsiveness | For agents and latency-sensitive workloads, versus “the nearest alternative” |
Nvidia frames both figures around responsiveness: how fast tokens come back to an agent or a user, rather than how many tokens a system produces in total.
How it fits Vera Rubin
The LPX is part of the Vera Rubin ecosystem, which Nvidia describes as “seven chips and five purpose-built racks.” According to the release, the LPX extends the Vera Rubin NVL72 platform and works together with Nvidia BlueField-4 DPUs and Nvidia Spectrum-6 SPX Ethernet. Nvidia is positioning it as the latency-focused member of the Vera Rubin family rather than as a standalone product line.
Who is deploying it
Nvidia names two early adopters:
- Nebius, the first AI cloud adopter, which is bringing the LPX to its Nebius Token Factory service.
- Groq, whose inference cloud platform is “among the first to bring” the system to market.
Context
The LPX carries the Groq name inside Nvidia’s own product line, and Groq’s inference cloud is among its first users. Two days after the LPX announcement, Nvidia reported second-quarter fiscal 2027 results on August 26, 2026, saying Vera Rubin “is ramping into full production with racks running at partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius.” Nebius therefore appears both as an early Vera Rubin rack operator and as the first LPX cloud.
What it changes
With the LPX in full production, Nvidia’s Vera Rubin generation includes a dedicated part for fast, latency-sensitive inference alongside its GPU racks, and two inference clouds have said they will offer it. The LPX and other Vera Rubin parts are tracked in our AI accelerator catalog.




