BTC $83,117.91 -0.37%
ETH $2,663.85 +0.25%
BNB $757.73 -1.84%
XRP $1.49 -0.71%
SOL $117.65 -1.95%
TRX $0.3341 +0.03%
DOGE $0.0931 -1.19%
ADA $0.2432 -2.23%
BCH $305.39 -3.53%
LINK $14.75 +6.03%
HYPE $87.34 -3.05%
AAVE $149.10 -0.67%
SUI $1.11 -9.17%
XLM $0.2258 +6.77%
ZEC $1,379.49 -11.86%
AAPL $337.78 -0.73%
AMZN $246.05 -1.15%
GOOGL $341.97 -0.09%
MSFT $507.41 -1.90%
META $712.04 -2.86%
NVDA $228.42 +1.96%
TSLA $356.84 -3.67%
SNDK $1,696.77 -2.27%
INTC $113.99 -4.44%
SPCX $145.82 -2.13%
MU $1,050.83 -1.51%
AMD $604.54 -2.19%
BTC $83,117.91 -0.37%
ETH $2,663.85 +0.25%
BNB $757.73 -1.84%
XRP $1.49 -0.71%
SOL $117.65 -1.95%
TRX $0.3341 +0.03%
DOGE $0.0931 -1.19%
ADA $0.2432 -2.23%
BCH $305.39 -3.53%
LINK $14.75 +6.03%
HYPE $87.34 -3.05%
AAVE $149.10 -0.67%
SUI $1.11 -9.17%
XLM $0.2258 +6.77%
ZEC $1,379.49 -11.86%
AAPL $337.78 -0.73%
AMZN $246.05 -1.15%
GOOGL $341.97 -0.09%
MSFT $507.41 -1.90%
META $712.04 -2.86%
NVDA $228.42 +1.96%
TSLA $356.84 -3.67%
SNDK $1,696.77 -2.27%
INTC $113.99 -4.44%
SPCX $145.82 -2.13%
MU $1,050.83 -1.51%
AMD $604.54 -2.19%

cere

All
Article
Flash

hot_img Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.
app_icon
ChainCatcher Building the Web3 world with innovations.