基础设施 4.0 · 优秀 2026-08-20 · 文章

Cerebras CS-4: rack-scale wafer inference, up to 30x faster than GPUs

Cerebras 发布 CS-4 机柜级推理系统:每套三片 WSE-3 Turbo 晶圆(官称单瓦较上代提速至 2 倍);宣称较 GPU 系统推理快至 30 倍,瓦间互连延迟降至 2 微秒,在超过 10 万亿参数的模型上保持 1000+ tokens/s;每瓦吞吐量较 CS-3 高至 10 倍;配 Nexus 机柜平台面向超大规模部署

打开原文回到归档

Cerebras product page for the newly announced CS-4: a rack-scale solution with three WSE-3 Turbo wafers per system, each wafer claimed up to 2x the speed of the previous generation. Headline claims: up to 30x faster inference than GPU systems; wafer-to-wafer interconnect latency reduced to 2 microseconds, sustaining 1,000+ tokens/sec on models exceeding 10 trillion parameters while preserving interactive decode; up to 10x more throughput per watt than CS-3; Nexus rack-scale platform for hyperscale datacenter deployment. All figures are vendor claims from the launch page; independent benchmarks not yet available at scan time.

Source: https://www.cerebras.ai/cs4