Analysis
The first benchmarks for Nvidia's Groq 3 LPU inference hardware have been published, offering an early read on a roughly $20 billion bet, The Register reported. Separately, Nvidia announced new customers for its Vera CPU and Groq LPX racks, The Information reported.
Groq's original architecture, developed by a team led by former Google TPU engineer Jonathan Ross, took a deliberately different approach from GPUs: a deterministic, software-scheduled design with on-chip SRAM instead of external memory, optimized for streaming tokens out fast rather than crunching training batches. Its selling point was always latency -- responses that arrive at conversational speed -- rather than throughput per dollar on a training job. Pulse has previously covered Groq's rise as an inference-specialist challenger before its technology and roughly $20 billion in commitments were absorbed into Nvidia's roadmap.
โIts selling point was always latency -- responses that arrive at conversational speed -- rather than throughput per dollar on a training job.โ
Groq raised at a multibillion-dollar valuation as an independent company on the strength of public demonstrations showing its LPUs generating tokens noticeably faster than GPU-based serving stacks, particularly for open-weight models like Llama and Mixtral running at high concurrency. That speed advantage is what made it an acquisition target rather than merely a partner: Nvidia's GPUs remain the default for training and for throughput-optimized inference, but a growing share of consumer-facing products -- voice agents, coding assistants, real-time translation -- are latency-bound rather than throughput-bound, and that is precisely the segment Groq's architecture was built to win.
Folding that into Nvidia's product line addresses a genuine strategic exposure. Training demand is concentrated among a handful of labs; inference demand scales with every deployed application, and it is the workload most likely to migrate to cheaper silicon from AMD, Google, Amazon or a specialist. A dedicated inference part inside the Nvidia stack means a customer optimizing inference cost does not have to leave the ecosystem to do it.
Benchmarks published this early are worth reading skeptically. Vendor-adjacent configurations, batch sizes and model choices all move inference numbers substantially, and the figure that decides purchasing is cost per million tokens at production concurrency, under a real serving stack. New named customers for the LPX racks are the more durable evidence, because those represent someone spending money rather than someone publishing a chart.