BTC $82,922.23 -0.73%
ETH $2,654.50 +0.12%
BNB $754.56 -2.46%
XRP $1.47 -2.25%
SOL $116.47 -3.31%
TRX $0.3345 +0.24%
DOGE $0.0918 -3.34%
ADA $0.2398 -5.34%
BCH $303.75 -6.83%
LINK $15.02 +7.14%
HYPE $86.01 -5.17%
AAVE $146.25 -3.26%
SUI $1.10 -11.82%
XLM $0.2230 +3.94%
ZEC $1,374.43 -12.82%
AAPL $338.00 -0.74%
AMZN $246.30 -0.98%
GOOGL $342.43 +0.17%
MSFT $508.75 -1.55%
META $716.44 -2.43%
NVDA $228.74 +1.99%
TSLA $357.24 -3.48%
SNDK $1,700.19 -2.09%
INTC $114.58 -4.30%
SPCX $145.82 -2.15%
MU $1,050.64 -1.57%
AMD $606.53 -2.27%
BTC $82,922.23 -0.73%
ETH $2,654.50 +0.12%
BNB $754.56 -2.46%
XRP $1.47 -2.25%
SOL $116.47 -3.31%
TRX $0.3345 +0.24%
DOGE $0.0918 -3.34%
ADA $0.2398 -5.34%
BCH $303.75 -6.83%
LINK $15.02 +7.14%
HYPE $86.01 -5.17%
AAVE $146.25 -3.26%
SUI $1.10 -11.82%
XLM $0.2230 +3.94%
ZEC $1,374.43 -12.82%
AAPL $338.00 -0.74%
AMZN $246.30 -0.98%
GOOGL $342.43 +0.17%
MSFT $508.75 -1.55%
META $716.44 -2.43%
NVDA $228.74 +1.99%
TSLA $357.24 -3.48%
SNDK $1,700.19 -2.09%
INTC $114.58 -4.30%
SPCX $145.82 -2.15%
MU $1,050.64 -1.57%
AMD $606.53 -2.27%

cere

All
Article
Flash

hot_img Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.
app_icon
ChainCatcher Building the Web3 world with innovations.