VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Nvidia Releases Nemotron 3.5 Lightning, Its First Open Model Since Huang's
Value Add VC/Pulse/AIDEEP DIVE

Nvidia Releases Nemotron 3.5 Lightning, Its First Open Model Since Huang's

Nvidia released Nemotron 3.5 Lightning, a 30-billion-parameter open-weight model free for commercial use, its first open release since CEO Jensen Huang publicly campaigned for open-weight AI policy in Washington last month.

By the Numbers

30B total
Parameters
3B (MoE)
Active per token
up to 4x faster output
Speed gain
~30% faster
Agentic task speed
150+ orgs
Open letter signatories
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
August 12, 2026
3 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Nvidia released Nemotron 3.5 Lightning on August 11, a 30-billion-parameter mixture-of-experts model with only 3 billion active parameters per token, distilled from the larger Nemotron 3 Ultra

2

The model is free for commercial use -- companies can download, modify and deploy it without seeking Nvidia's permission -- and runs on a single consumer GPU, including RTX PCs and Nvidia's DGX Spark systems

3

Nvidia says the model delivers up to 4x faster output and completes agentic tasks roughly 30% faster than comparable open models in its weight class

4

It's Nvidia's first open-weight model release since Jensen Huang's July 24 open letter, 'Open Weights and American AI Leadership,' urging Washington not to restrict open-weight AI -- a letter that launched with 25 signatories and grew past 150 organizations within days

TC

The VC Read · Trace's Take

Trace Cohen

The tell here isn't the benchmark numbers, it's the timing -- Nvidia turning Jensen Huang's open-weights policy letter into a shipped product three weeks later means the company is betting its hardware moat holds even when the software layer is free. The diligence question for anyone building on open-weight infrastructure: does Lightning's speed advantage hold up once developers move past Nvidia's own benchmarks, and does 'free' model weights still mean 'locked into Nvidia GPUs' in practice? Watch whether major agent frameworks adopt it natively in the next month -- that's the real adoption signal, not the release headline.

AI Chip Wars →

Analysis

The Release

Nvidia released Nemotron 3.5 Lightning on August 11, a 30-billion-parameter open-weight model built on a mixture-of-experts architecture that activates only 3 billion parameters per token, according to CNBC. The model is distilled from Nvidia's larger Nemotron 3 Ultra and is designed specifically for high-volume agentic AI workloads rather than general-purpose chat. It's free for commercial use -- companies can download, use and modify it without seeking permission from Nvidia. Nvidia's own release notes, published alongside a companion routing tool called NeMo Switchyard, detail the same specs and describe the pairing as built to send each agent request to the most efficient model for the task, according to the Nvidia Blog.

What's Different About It

Because Nemotron 3.5 Lightning activates just 3 billion of its 30 billion parameters per token, it can run on a single consumer GPU -- an RTX-equipped PC or Nvidia's DGX Spark desktop system -- rather than requiring the data-center-scale hardware most frontier models need. Nvidia says the model delivers up to 4x faster output speed and completes agentic tasks roughly 30% faster than comparable open models in its class, a claim aimed squarely at developers building autonomous agents that need to run many fast inference calls rather than a single high-quality response.

Why Nvidia Is Doing This Now

The release lands three weeks after Jensen Huang published his first post on X on July 24: an industry letter titled "Open Weights and American AI Leadership," arguing Washington should avoid restricting open-weight AI models. The letter launched with 25 signatories and grew to more than 150 organizations within days. Nemotron 3.5 Lightning is Nvidia's first open-weight model since that campaign, turning a policy argument into a concrete product release -- a company that sells the chips nearly every AI lab depends on now also has a direct commercial stake in whether open-weight models remain legal to build and distribute freely.

The Competitive Field

Nvidia's open-weight push puts it in more direct competition with Meta's Llama family, Mistral, and Chinese open-weight labs like Alibaba's Qwen and DeepSeek, all of which have used free, modifiable model weights to build developer mindshare that a closed API can't match. It's a different posture than Nvidia's core hardware business, where the company profits regardless of which model developers run -- but a widely adopted Nvidia open model could steer more of that inference workload toward Nvidia's own optimized software stack and hardware, the same way CoreWeave's record quarter this week showed how much demand is flowing through Nvidia-dependent infrastructure already.

Numbers in Context

A 30-billion-parameter model that runs on a single consumer GPU is small by frontier-lab standards -- OpenAI, Anthropic and Google's flagship models are believed to run into the hundreds of billions or trillions of parameters on data-center clusters -- but that's the point: Lightning is built for cheap, high-volume agentic inference, not for competing on raw capability with GPT-5-class models. It's a complementary product to Nvidia's $500 billion AI financing push this month, which is aimed at the opposite end of the market -- financing the data-center-scale compute that frontier labs need.

The Counterweight

Nvidia giving away a capable open model for free is also a hardware-sales strategy, not pure altruism -- widespread adoption of a Lightning-optimized workflow makes Nvidia GPUs the default choice for running it, even though the weights themselves are open. And a model tuned for speed and low parameter-count activation is a different bet than the raw-capability race dominating headlines elsewhere this week; Lightning's real test is whether developers building production agents choose it over Llama or Qwen, not whether it tops a capability leaderboard.

Ahead

Watch adoption metrics -- downloads, fine-tunes, and whether major agent frameworks add native Lightning support -- over the next month; that's the signal for whether Nvidia's open-weight bet converts into the kind of developer lock-in effect Meta's Llama achieved, or remains a smaller niche release next to the frontier labs' closed flagship models.

ShareXLinkedInEmail

More on

Nvidia →

Reported by CNBC · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 12, 2026

CoreWeave Revenue More Than Doubles, Stock Jumps 18%

Illustration for: CoreWeave Revenue More Than Doubles, Stock Jumps 18%
AI

CoreWeave Revenue More Than Doubles, Stock Jumps 18%

CoreWeave reported second-quarter revenue of $2.6 billion, up 112% year-over-year, and raised its full-year guidance on a backlog that has swelled past $104 billion, sending shares up 18% in premarket trading.

AI· Aug 11, 2026

Gemini Hits 1 Billion Monthly Users, Catching Up to ChatGPT

Illustration for: Gemini Hits 1 Billion Monthly Users, Catching Up to ChatGPT
AI

Gemini Hits 1 Billion Monthly Users, Catching Up to ChatGPT

Google's Gemini app crossed 1 billion monthly users, its fastest-growing product ever, closing ground on ChatGPT even as OpenAI reports its own product separately hit 1 billion weekly users on a different growth curve.

AI· Aug 11, 2026

Decagon Crosses $100M ARR Betting Against Forward-Deployed Engineers

Illustration for: Decagon Crosses $100M ARR Betting Against Forward-Deployed Engineers
AI

Decagon Crosses $100M ARR Betting Against Forward-Deployed Engineers

Decagon, a three-year-old AI customer-service agent startup valued at $4.5 billion, crossed $100 million in annualized revenue while deliberately avoiding the forward-deployed engineer model rivals like Sierra use to customize deployments.

@Trace_Cohen·t@nyvp.com