The latest AI Beat newsletter reports that NVIDIA's Groq 3 LPX has entered full-scale production. In December of last year, NVIDIA spent around $20 billion to acquire Groq's technology, poaching founder Jonathan Ross and other key team members. Eight months later, the first hardware from this deal has been officially rolled out.
The Groq 3 LPX is a system designed specifically to accelerate large-scale model inference, incorporating 256 Groq 3 LPU chips. In Artificial Analysis testing, when Gemma 4 with a 31B input ran 100,000 Tokens, the Groq 3 LPX processed 3,431 output Tokens per second. In the same test, the fastest publicly available API at the time handled only about 870 Tokens, making the Groq 3 LPX nearly 4 times faster.
In actual deployment, the Rubin GPU handles heavier computations, while the Groq 3 LPX is dedicated to speeding up Token generation. For Coding Agent, when calling the model dozens or hundreds of times consecutively, there is a noticeable reduction in the model's "speaking" time.
Nebius will be the first AI cloud provider to adopt the Groq 3 LPX. Interestingly, Groq itself has become one of the early adopters and will also collaborate with Dell on deployment. After licensing the technology to NVIDIA last year, this year they are using NVIDIA's own "Groq" technology.

