NVIDIA announces full-scale production of Groq racks, which will be deployed at Nebius
On Monday local time, Nvidia Senior Director Dion Harris announced that the Groq 3 LPX rack has entered full-scale production and will be deployed in the data center of the new cloud service provider Nebius, expected to go live later this year. This move marks the commercialization of the technology obtained by Nvidia after reaching a technology licensing deal worth approximately $20 billion with Groq last December. Several core employees from Groq have joined Nvidia, with founder Jonathan Ross serving as Nvidia's Chief Software Architect.
Nvidia is accelerating the production of Groq chips and providing products to customers, highlighting the importance of low-latency inference. The Groq 3 LPX, released in March this year, is an inference accelerator for the Vera Rubin platform, integrating 500 megabytes of high-speed SRAM on the chip die to reduce memory bottlenecks. Each LPX rack can integrate 256 Groq 3 chips, and Nvidia cites benchmarks indicating it can process 3,400 tokens per second; the chip is manufactured by Samsung.
Harris stated that low-latency chips are not meant to replace GPUs but are designed to use the appropriate processor for different stages of workloads, allowing cloud service providers to offer higher-priced packages for latency-sensitive users. Nvidia CEO Jensen Huang plans to allocate about a quarter of the data center space for programming applications to Groq chips, with the remainder using the Vera Rubin system, and expects cumulative sales of Blackwell and Vera Rubin to reach $1 trillion by 2027.






