Illustration for: First Groq LPU Benchmarks Test Nvidia's $20B Bet

First Groq LPU Benchmarks Test Nvidia's $20B Bet

Early benchmarks for Nvidia's Groq-derived LPU inference hardware are out, alongside new customer announcements for its Vera CPU and Groq LPX racks.

TC
By the AI Desk
Edited by Trace Cohen ยท Early-stage VC & angel ยท Founder, New York Venture Partners
Updated August 25, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The first Groq 3 LPU benchmarks give a read on Nvidia's roughly $20 billion inference bet, [The Register reported](https://www.theregister.com/systems/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-tell-us-about-its-20b-gamble/5291880)

2

Nvidia has separately announced new customers for its Vera CPU and Groq LPX racks, [per The Information](https://www.theinformation.com/briefings/nvidia-announces-new-customers-vera-cpu-groq-lpx-racks)

3

Inference is the larger long-run market, and it rewards different silicon than training does -- latency and cost per token rather than raw FLOPs

4

Owning both the training GPU and a dedicated inference part is how Nvidia defends against customers moving inference to cheaper alternatives

TC

The VC Read ยท Trace's Take

Trace Cohen

Inference-specialist startups just got a harder pitch: their differentiation was latency, and the incumbent now sells latency too. If you hold a position in one, the question for the next board meeting is which workloads still justify a non-Nvidia part after this, and whether any customer will sign a multi-year commitment for them.

Analysis

The first benchmarks for Nvidia's Groq 3 LPU inference hardware have been published, offering an early read on a roughly $20 billion bet, The Register reported. Separately, Nvidia announced new customers for its Vera CPU and Groq LPX racks, The Information reported.

Groq's original architecture, developed by a team led by former Google TPU engineer Jonathan Ross, took a deliberately different approach from GPUs: a deterministic, software-scheduled design with on-chip SRAM instead of external memory, optimized for streaming tokens out fast rather than crunching training batches. Its selling point was always latency -- responses that arrive at conversational speed -- rather than throughput per dollar on a training job.

โ€œUpdate (August 25, 2026): Pulse has follow-up coverage โ€” Nvidia Announces New Vera CPU, Groq LPX Rack Customers.โ€

Folding that into Nvidia's product line addresses a genuine strategic exposure. Training demand is concentrated among a handful of labs; inference demand scales with every deployed application, and it is the workload most likely to migrate to cheaper silicon from AMD, Google, Amazon or a specialist. A dedicated inference part inside the Nvidia stack means a customer optimizing inference cost does not have to leave the ecosystem to do it.

Benchmarks published this early are worth reading skeptically. Vendor-adjacent configurations, batch sizes and model choices all move inference numbers substantially, and the figure that decides purchasing is cost per million tokens at production concurrency, under a real serving stack. New named customers for the LPX racks are the more durable evidence, because those represent someone spending money rather than someone publishing a chart.

Update (August 25, 2026): Pulse has follow-up coverage โ€” Nvidia Announces New Vera CPU, Groq LPX Rack Customers.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by The Register ยท Analysis by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.