Cerebras Launches CS-4: A Multi-Wafer System for AI Inference
Cerebras has introduced the CS-4, its first multi-wafer system, which packs three of its dinner-plate-sized processors into a single rack. Unveiled on Tuesday at the company’s Supernova event, the CS-4 promises to run frontier models up to 30 times faster than GPU-based systems.
Key Features:
- Performance: Carries three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and supports models exceeding 50 trillion parameters.
- Latency: Reduces wafer-to-wafer latency to two microseconds from five.
- Power Efficiency: Uses half as many components as its predecessor and moves power conversion closer to the processors, resulting in estimated power consumption of 120-140 kilowatts per rack, roughly half that of comparable AMD and Nvidia rack systems.
Debating Newness:
While the CS-4 boasts impressive specifications, it’s important to note that the chip inside is not entirely new. The WSE-3 Turbo retains the same four trillion transistors, core count, and manufacturing process as its predecessor, with the main enhancements being a clock bump rather than a redesign.
Target Market:
Cerebras targets enterprise buyers, emphasizing the CS-4’s speed and power efficiency for AI applications beyond chatbots, such as reasoning, verification, and tool use.
Partnerships:
The company announced partnerships with OpenAI, G42, MBZUAI, and AWS, and included AMD’s Helios rack in its list of partner systems, reflecting its commitment to working with all major players in AI hardware except Nvidia.
Financial Outlook:
Cerebras reported second-quarter revenue of $180.1 million, a 74% year-on-year increase, driven primarily by cloud revenue. However, sequential revenue declined from $193.4 million in Q1, and margin pressure persists. The company incurred a GAAP net loss of $450.5 million and an adjusted loss of $6.9 million for the quarter.