Illustration for: Perplexity, Nvidia Launch Local AI Agent, Zero Tokens

Perplexity, Nvidia Launch Local AI Agent, Zero Tokens

Perplexity and Nvidia launched Portable Computer, a fully local version of Perplexity's agent platform running on RTX GPUs or Nvidia's DGX Spark, eliminating per-token API costs for users who can supply their own hardware.

By the Numbers

RTX, 24GB VRAM
Min. GPU requirement
DGX Spark, RTX Linux
Supported hardware
Linux (Aug 2026)
Launch platform
Sept 2026
Windows support
$0
Token cost locally
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
Updated August 26, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Perplexity and Nvidia launched Portable Computer, a fully local version of Perplexity's agentic Computer platform that runs entirely on user hardware with no cloud API billing, [VentureBeat reported](https://venturebeat.com/infrastructure/perplexity-partners-with-nvidia-to-launch-portable-computer-a-fully-local-ai-agent-with-zero-token-costs)

2

It requires an RTX GPU with at least 24GB of VRAM or Nvidia's DGX Spark desktop supercomputer, launching on Linux this month with Windows support planned for September; Apple silicon is absent from the roadmap

3

The product competes directly with Ollama on local inference, but Perplexity is positioning its differentiation at the agent-orchestration layer rather than raw model-serving, citing internal benchmarks against open-source harnesses

4

For any startup building a metered, API-billed AI agent product, a credible zero-marginal-cost local alternative from two well-capitalized players is a real pricing-pressure signal worth modeling into unit economics

TC

The VC Read · Trace's Take

Trace Cohen

For any founder building a metered AI-agent product priced per API call, this is the competitive pressure to model now, not later -- a credible local alternative from Nvidia and Perplexity puts a real ceiling on what heavy users will tolerate paying for equivalent cloud usage. The hardware barrier (24GB VRAM minimum) is the actual moat protecting cloud-billed products for now, but that bar drops every GPU generation, and I'd want any agent-product pitch to have a plan for when local inference gets cheap enough to matter to their core user base.

Analysis

Perplexity and Nvidia jointly launched Portable Computer, a fully local version of Perplexity's agentic Computer platform that runs entirely on a user's own hardware rather than cloud infrastructure, VentureBeat reported. The product bundles models, an inference engine, tool connectors, security sandboxing and the agent-orchestration layer into a single application, extending a partnership between the two companies that began in June 2025 around sovereign AI initiatives.

The hardware requirement is real but not exotic: an RTX GPU with at least 24GB of VRAM -- roughly a GeForce RTX 3090 or newer -- or Nvidia's DGX Spark desktop supercomputer, with multiple Sparks able to link over shared memory for larger models. The product launched Linux-only this month, with Windows support arriving in September; Apple silicon support is not on the current roadmap, a meaningful gap given how much AI-development and prosumer usage happens on Mac hardware.

  • Perplexity -- providing the agent-orchestration software layer (Computer) and its custom PPLX 27B model variant
  • Nvidia -- providing the local-hardware target (DGX Spark, RTX GPUs) and its Nemotron 3.5 Lightning model, coming soon to the platform
  • Ollama -- the closest existing competitor on local inference, though VentureBeat's reporting frames the distinction as Ollama solving the inference layer while Portable Computer targets the full agent-harness layer on top of it
  • Qwen 3.8 27B -- one of the supported open models, alongside Perplexity's own PPLX variant

Update (August 26, 2026): Pulse has follow-up coverage — Apple's New Mac Studio, Mini Lean Hard Into Local AI.

"Zero token costs" specifically means inference running locally doesn't accrue the per-call API billing that makes extended agentic tasks -- document review, verification loops, iterative multi-step analysis -- expensive to run against cloud models. Perplexity's own testing reportedly found that models like Qwen struggle with reliable performance past roughly 100,000 tokens of context despite being marketed with 260,000-token windows, which is part of the rationale for a purpose-built local agent harness rather than a general-purpose framework layered on top of raw inference.

The honest limitation is that local execution shifts cost from per-token API billing to upfront hardware cost -- a 24GB-VRAM GPU or DGX Spark is not a trivial purchase for an individual user, and "zero token costs" only looks attractive relative to sustained heavy cloud-API usage, not against a user's total cost of ownership including the hardware itself. Perplexity's benchmarks claiming advantages over open-source harnesses like Pi and Hermes also remain internal and haven't been independently validated by third parties.

Users who retain the option to escalate specific complex tasks to Perplexity's frontier cloud models, with explicit permission and cost transparency, get a hybrid model that hedges the tradeoff -- local-first for routine tasks, cloud on demand for anything the local models can't handle reliably, which is closer to how the product will likely get used in practice than either pure-local or pure-cloud framing suggests.

Update (August 26, 2026): Pulse has follow-up coverage — Apple's New Mac Studio, Mini Lean Hard Into Local AI.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.