VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog๐ŸคPartner
Illustration for: First Groq LPU Benchmarks Test Nvidia's $20B Bet
Value Add VC/Pulse/AIDEEP DIVE$20B

First Groq LPU Benchmarks Test Nvidia's $20B Bet

Early benchmarks for Nvidia's Groq-derived LPU inference hardware are out, alongside new customer announcements for its Vera CPU and Groq LPX racks.

NvidiaGroq
TC
By the AI Desk
Edited by Trace Cohen ยท Early-stage VC & angel ยท Founder, New York Venture Partners
August 24, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The first Groq 3 LPU benchmarks give a read on Nvidia's roughly $20 billion inference bet, [The Register reported](https://www.theregister.com/systems/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-tell-us-about-its-20b-gamble/5291880)

2

Nvidia has separately announced new customers for its Vera CPU and Groq LPX racks, [per The Information](https://www.theinformation.com/briefings/nvidia-announces-new-customers-vera-cpu-groq-lpx-racks)

3

Inference is the larger long-run market, and it rewards different silicon than training does -- latency and cost per token rather than raw FLOPs

4

Owning both the training GPU and a dedicated inference part is how Nvidia defends against customers moving inference to cheaper alternatives

TC

The VC Read ยท Trace's Take

Trace Cohen

Inference-specialist startups just got a harder pitch: their differentiation was latency, and the incumbent now sells latency too. If you hold a position in one, the question for the next board meeting is which workloads still justify a non-Nvidia part after this, and whether any customer will sign a multi-year commitment for them.

AI Chip Startups โ†’ AI Chip Wars โ†’

Analysis

The first benchmarks for Nvidia's Groq 3 LPU inference hardware have been published, offering an early read on a roughly $20 billion bet, The Register reported. Separately, Nvidia announced new customers for its Vera CPU and Groq LPX racks, The Information reported.

Groq's original architecture, developed by a team led by former Google TPU engineer Jonathan Ross, took a deliberately different approach from GPUs: a deterministic, software-scheduled design with on-chip SRAM instead of external memory, optimized for streaming tokens out fast rather than crunching training batches. Its selling point was always latency -- responses that arrive at conversational speed -- rather than throughput per dollar on a training job. Pulse has previously covered Groq's rise as an inference-specialist challenger before its technology and roughly $20 billion in commitments were absorbed into Nvidia's roadmap.

โ€œIts selling point was always latency -- responses that arrive at conversational speed -- rather than throughput per dollar on a training job.โ€

Groq raised at a multibillion-dollar valuation as an independent company on the strength of public demonstrations showing its LPUs generating tokens noticeably faster than GPU-based serving stacks, particularly for open-weight models like Llama and Mixtral running at high concurrency. That speed advantage is what made it an acquisition target rather than merely a partner: Nvidia's GPUs remain the default for training and for throughput-optimized inference, but a growing share of consumer-facing products -- voice agents, coding assistants, real-time translation -- are latency-bound rather than throughput-bound, and that is precisely the segment Groq's architecture was built to win.

Folding that into Nvidia's product line addresses a genuine strategic exposure. Training demand is concentrated among a handful of labs; inference demand scales with every deployed application, and it is the workload most likely to migrate to cheaper silicon from AMD, Google, Amazon or a specialist. A dedicated inference part inside the Nvidia stack means a customer optimizing inference cost does not have to leave the ecosystem to do it.

Benchmarks published this early are worth reading skeptically. Vendor-adjacent configurations, batch sizes and model choices all move inference numbers substantially, and the figure that decides purchasing is cost per million tokens at production concurrency, under a real serving stack. New named customers for the LPX racks are the more durable evidence, because those represent someone spending money rather than someone publishing a chart.

Related Deep Dives

  • How Does Groq Make Money: LPU Chips, GroqCloud Tokens, an... โ†’
  • AI Product Costs โ€” GPU, API & Inference (2026) โ†’
  • 3B+ Downloads โ€” Qwen Open-Weight Model Rankings โ†’
ShareXLinkedInEmail

More on

Nvidia โ†’Groq โ†’

Prior Pulse Coverage

NvidiaNvidia Weighs Perplexity Investment Above $30BNvidiaNvidia to Raise Flagship AI Chip Prices 17%NvidiaThe Gaps in Nvidia's $500B Financing PitchNvidiaNvidia and Salesforce Earnings Take the SpotlightNvidiaNvidia's Cloverleaf Deal Is Patching AI Bubble Cracks

Key Sources

2 sources
SourceThe Register
AnalysisValue Add Pulse

Reported by The Register ยท Analysis by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AIยท Aug 24, 2026

Hugging Face in Talks to Sell for $13 Billion

Illustration for: Hugging Face in Talks to Sell for $13 Billion
AI$13B

Hugging Face in Talks to Sell for $13 Billion

Hugging Face, the New York-based host of the open-source AI model ecosystem, is reportedly in talks to be acquired in a deal valuing it around $13 billion, roughly triple its last private mark.

AIยท Aug 24, 2026

Nvidia Weighs Perplexity Investment Above $30B

Illustration for: Nvidia Weighs Perplexity Investment Above $30B
AI$30B+ valuation

Nvidia Weighs Perplexity Investment Above $30B

Nvidia has discussed investing in Perplexity at a valuation above $30 billion and separately considered a technology licensing arrangement with the AI search company.

AIยท Aug 24, 2026

The Gaps in Nvidia's $500B Financing Pitch

Illustration for: The Gaps in Nvidia's $500B Financing Pitch
AI$500B

The Gaps in Nvidia's $500B Financing Pitch

A closer look at Nvidia's $500 billion financing pitch finds the sources of capital behind the headline figure are considerably less settled than the number implies.

Deep Dives

How Does Groq Make Money: LPU Chips, GroqCloud Tokens, an...AI Product Costs โ€” GPU, API & Inference (2026)3B+ Downloads โ€” Qwen Open-Weight Model Rankings
@Trace_Cohenยทt@nyvp.com