BTC $76,975.78 -0.38%
ETH $2,483.46 -1.67%
BNB $716.27 -1.52%
XRP $1.35 -1.45%
SOL $99.64 -2.07%
TRX $0.3379 -0.60%
DOGE $0.0826 -2.58%
ADA $0.2044 -1.64%
BCH $221.49 -2.16%
LINK $11.26 -2.13%
HYPE $77.62 -2.39%
AAVE $125.05 -0.66%
SUI $0.7030 -3.03%
XLM $0.1782 -1.02%
ZEC $1,066.10 -5.44%
AAPL $330.51 -0.74%
AMZN $255.07 -0.72%
GOOGL $337.18 -1.00%
MSFT $492.83 -0.60%
META $642.99 -0.86%
NVDA $215.27 -1.50%
TSLA $361.10 -1.75%
SNDK $1,561.87 -4.21%
INTC $99.25 -2.62%
SPCX $148.88 -0.86%
MU $941.79 -2.73%
AMD $503.80 -2.26%
BTC $76,975.78 -0.38%
ETH $2,483.46 -1.67%
BNB $716.27 -1.52%
XRP $1.35 -1.45%
SOL $99.64 -2.07%
TRX $0.3379 -0.60%
DOGE $0.0826 -2.58%
ADA $0.2044 -1.64%
BCH $221.49 -2.16%
LINK $11.26 -2.13%
HYPE $77.62 -2.39%
AAVE $125.05 -0.66%
SUI $0.7030 -3.03%
XLM $0.1782 -1.02%
ZEC $1,066.10 -5.44%
AAPL $330.51 -0.74%
AMZN $255.07 -0.72%
GOOGL $337.18 -1.00%
MSFT $492.83 -0.60%
META $642.99 -0.86%
NVDA $215.27 -1.50%
TSLA $361.10 -1.75%
SNDK $1,561.87 -4.21%
INTC $99.25 -2.62%
SPCX $148.88 -0.86%
MU $941.79 -2.73%
AMD $503.80 -2.26%
hot_img

Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

2026-08-19 10:38:29

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.

Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.

app_icon
ChainCatcher Building the Web3 world with innovations.