Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
Nvidia is moving its Groq 3 LPX inference chip into full production and reports 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras. But the numbers don't tell the whole story. Nvidia needs at least 64 accelerators to get there, while Cerebras needs only one or two, according to…
Annons
Annons
Source: The Decoder — Published — Category: Models
More from The Decoder today
Anthropic's Claude can now orchestrate up to 1,000 AI agents in parallel through dynamic workflows 11h ago Anthropic launches a free AI scanner for open-source projects 11h ago OpenAI revenue keeps surging as company seeks $30 billion in fresh capital 12h ago
Annons
Annons