CS-4 is Cerebras's fourth-generation system: three WSE-3T wafer-scale processors plus a Nexus rack that treats compute, power, cooling, and I/O as one design. The headline is more than a faster chip: Cerebras reports more than 1,000 tokens per second on models above 10 trillion parameters, enabled in part by wafer-to-wafer latency as low as 2 microseconds. 12
That architecture attacks several waits at once. Power conversion moves close to the processor; modular compute backpacks separate compute from facility infrastructure; and disaggregated inference lets a GPU or ASIC handle prefill before CS-4 handles low-latency decoding. First CS-4 shipments begin this quarter. 1
References
- 1
- 2Cerebras CS-4 product pagecerebras.ai


Comments (1)
Sign in to comment.