Cerebras Systems unveiled the CS-4 on August 18, 2026, a system built on its new WSE-3 Turbo (WSE-3T) wafer-scale processor, according to the company’s announcement. In the August 18 release, Cerebras said “first CS-4 shipments begin this quarter.”

On this page

Specifications Cerebras published

Each WSE-3T is a single wafer-sized processor. A CS-4 rack holds three of them.

Measure Per WSE-3T wafer Per CS-4 rack (3 wafers)
Transistors 4 trillion —
AI cores 900,000 —
On-wafer SRAM 44GB —
AI compute 250 PFLOPS 750 PFLOPS
Memory bandwidth 43.2 PB/s 129.6 PB/s

The wafers connect through what Cerebras calls Direct Wafer Links, which it says connect wafers within and across racks without a switch and bring wafer-to-wafer latency “as low as two microseconds.”

Speed claims

All performance figures below are Cerebras’s own. The company says that “in a head-to-head comparison on GPT-OSS-120B, when given identical prompts, the CS-4 delivers more than 4,400 tokens second per user (TPS/user), up to 30 times faster than GPU solutions.”

Against its own previous system, Cerebras says the CS-4 is “up to twice as fast as the CS-3” and “delivers up to 10x more throughput per watt than the CS-3,” which it says improves data-center economics.

Context: the OpenAI Ultrafast mode

Five days earlier, on August 13, 2026, Cerebras said it powers Ultrafast mode, “a new service tier in the OpenAI API for GPT-5.6 Sol.” Cerebras says the tier runs at “up to 750 output tokens per second,” up to 14 times faster than OpenAI’s standard processing, and is “available initially in limited preview to OpenAI customers.”

What it changes

The CS-4 groups three wafers into one rack joined by Direct Wafer Links, and Cerebras measures the result in tokens per second per user: the speed a single user of a model sees, which is what interactive and agentic applications depend on. The OpenAI tier, announced the same month, is a commercial example of selling that kind of per-user speed through a major lab’s API. Wafer-scale and GPU-based accelerators are compared in our AI accelerator catalog.

Sources