Illustration for: Perplexity Open-Sources Its On-Device AI Engine

Perplexity Open-Sources Its On-Device AI Engine

Perplexity shipped hybrid compute for its Mac app and open-sourced Lily, a local inference engine, plus Privacy Gate, an on-device classifier that screens data before it reaches the cloud.

By the Numbers

1.23x faster
Lily prefill vs MLX-LM
1.35x faster
Lily decode vs MLX-LM
3
Local models at launch
Sept 1, 2026
Hybrid compute launched
Sept 2, 2026
Lily open-sourced
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Perplexity introduced hybrid compute for its Mac app on Sept. 1, splitting each Perplexity Computer task between frontier models in the cloud and a local model running on the Mac itself -- private files and sensitive actions stay on-device by default.

2

The next day, Perplexity open-sourced Lily, a local inference engine written in Rust with hand-written Metal kernels, claiming 1.23x higher prefill throughput and 1.35x higher decode throughput than Apple's own MLX-LM on the same model.

3

Alongside Lily, Perplexity open-sourced Privacy Gate on Hugging Face -- a small on-device classifier that inspects outbound data for PII (names, addresses, account numbers) and lets the user decide before anything sensitive reaches the cloud.

4

The releases land the same week Nvidia's $12.93 billion acquisition of Hugging Face raised questions about the platform's neutrality -- Perplexity publishing safety-relevant tooling there anyway is a practical vote of confidence in Hugging Face staying open regardless of who owns it.

TC

The VC Read · Trace's Take

Trace Cohen

Lily's 1.23x/1.35x throughput claims are Perplexity's own benchmark on Perplexity's own hardware choice -- worth independent verification before any portfolio company builds on top of it, since vendor-published inference benchmarks have a track record of not replicating cleanly on different workloads. The more interesting signal is Perplexity publishing Privacy Gate on Hugging Face the same week Nvidia's acquisition closed -- that's a real usage vote on Hub neutrality, worth more than any pledge in a press release.

Analysis

Perplexity introduced hybrid compute for its Mac app on Sept. 1, a mode that splits each Perplexity Computer task between frontier models running in the cloud and a local model running on the Mac itself, Unite.AI reported. Under the new architecture, the cloud handles frontier reasoning, web search and planning, while the local model processes private files, sensitive information and on-device actions -- keeping anything sensitive from leaving the machine by default rather than relying on a post-hoc redaction step.

The next day, Sept. 2, Perplexity open-sourced Lily, a local inference engine written in Rust with hand-written Metal kernels built specifically for Apple Silicon, MarkTechPost reported. On the Qwen3.6-35B-A3B model, Perplexity's own benchmarks show Lily running 1.23x faster on prefill throughput and 1.35x faster on decode throughput than Apple's own MLX-LM framework -- a direct challenge to Apple's on-device inference tooling, published by a company that doesn't make the hardware it's optimizing for.

Hybrid compute ships with three components:

- Three local models at launch -- Gemma 4 E4B, Qwen3.6 35B-A3B, and a Perplexity model post-trained specifically for Perplexity Computer.

  • Lily -- the local inference engine itself, open-sourced and benchmarked against MLX-LM.
  • Privacy Gate -- a small on-device classifier, open-sourced on Hugging Face, that inspects outbound data for PII such as names, addresses and account numbers and lets the user decide before anything sensitive is sent to the cloud.
  • Three local models at launch -- Gemma 4 E4B, Qwen3.6 35B-A3B, and a Perplexity model post-trained specifically for Perplexity Computer.

The releases land the same week Nvidia's $12.93 billion acquisition of Hugging Face raised questions about whether the platform stays neutral once a chipmaker owns it. Perplexity choosing to publish safety-relevant tooling like Privacy Gate on Hugging Face anyway, the same week that deal closed, is a practical vote of confidence that the hub remains usable regardless of ownership -- a real-world signal that matters more than any acquirer's public pledge about staying compute-agnostic.

For AI application developers, Lily's benchmark claims against MLX-LM are worth independent verification before building on them -- vendor-published throughput comparisons on a vendor's own benchmark are common this cycle, and Apple hasn't yet responded publicly. The more durable signal is architectural: on-device PII screening before cloud calls is becoming a genuine product category, not just a compliance checkbox, and any startup building AI products that touch sensitive user data should be evaluating whether a similar local-gate step belongs in its own pipeline.

ShareXLinkedInEmail

Key Sources

3 sources
SourceUnite.AI
SupportUnite.AI

Reported by Unite.AI · First reported by Unite.AI · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.