Cerebras Launches CS-4, a High-Performance AI Inference Solution
Cerebras has unveiled the CS-4, a new AI inference system that claims to be up to 30 times faster than traditional GPU systems. The design features significant innovations in power delivery, cooling, and modular architecture, allowing for rapid deployment in large-scale data centers.
The Cerebras CS-4 is a rack-scale AI inference solution that offers up to 30 times faster inference than GPU systems.
Each wafer in the CS-4 delivers double the speed of its predecessor, the CS-3, while achieving up to 10 times more throughput per watt.
The system reduces interconnect latency to 2 microseconds, enabling over 1,000 tokens per second for models with more than 10 trillion parameters.
CS-4 is built on the new Cerebras Nexus Platform Architecture, featuring a modular design that simplifies manufacturing, deployment, and maintenance.
The Wafer-Scale Backpack integrates the wafer, power conversion, cooling, and control electronics into a compact package, reducing deployment time significantly.
Power delivery is optimized, being only 0.5 millimeters from the processor, which minimizes power loss and allows for higher operating frequencies.
A new programmable I/O subsystem in CS-4 doubles I/O bandwidth and lowers latency, facilitating connections between wafers without the need for a switch.
The Cerebras PowerRack can be pre-installed, allowing compute backpacks to be added quickly, streamlining the deployment process in hyperscale environments.