Illustration for: Arm Bets Phones Will Run Agents, Not Just Apps

Arm Bets Phones Will Run Agents, Not Just Apps

Arm's next mobile chip platform, CSS for Mobile 2, is built to run AI agents and console-quality graphics locally on a phone, moving inference off the cloud and onto the device inside a 1-watt power budget.

By the Numbers

CSS for Mobile 2
Platform
70%
Small-model speedup
1W
Gaming power budget
70%
Ray-tracing workload cut
2027
Expected in devices
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
3 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

Every chipmaker is now building for a world where phones run continuous background agents, not occasional chat queries -- that's the real signal buried in a spec sheet. If you're building mobile AI product, the architecture bet to make now is split inference: local for the cheap steps, cloud for the hard ones. Arm's own numbers are unverified until Qualcomm or MediaTek ships silicon on this design, likely 2027 -- treat every figure here as a roadmap claim, not a benchmark.

Analysis

Arm introduced its next mobile chip platform this week, CSS for Mobile 2, built around two priorities that used to trade off against each other: running AI agents locally and delivering console-quality graphics, The Register reported Tuesday. The platform is a licensable design, meaning phone makers still have to fabricate their own silicon around it -- the earliest devices are expected in 2027. Pulse has previously covered Arm's push into AI-capable silicon as the company positions its licensable designs against custom hyperscaler chips.

What's actually new in the silicon

The flagship compute cluster pairs two new C2-Ultra CPU cores -- Arm claims a 15% single-thread gain over last year's C1-Ultra -- with six efficiency cores and two Scalable Matrix Extension units for AI math, doubled from the prior generation. Arm says that doubling delivers a 70% speedup running small language models directly on-device. The new Mali G2-Ultra NX GPU adds neural-assisted ray tracing and upscaling, claiming a 70% cut in ray-tracing workload through AI techniques, enough for 30 frames-per-second ray-traced gaming inside a 1-watt power budget -- a fraction of what a discrete GPU draws for comparable output. The first silicon built on that GPU is already shipping: Xiaomi's Xring O3 chip, launching in the Xiaomi 18 Fold, pairs the Mali G2-Ultra NX with Neural Super Sampling and claims an 85% GPU performance gain over its predecessor, The Verge reported the same day -- an early data point on whether Arm's lab numbers hold up in a shipping device.

Arm says that doubling delivers a 70% speedup running small language models directly on-device.

Why the target is agents, not chatbots

Every major AI lab has spent 2026 shipping agent products that call tools, browse the web and act across multiple steps without a human approving each one. Running that reliably on a phone today means either a fast round-trip to a cloud model -- with the latency, battery and privacy costs that implies -- or a heavily quantized local model that trades accuracy for speed. Arm's bet is that on-device inference becomes commercially necessary once agents are doing continuous background work rather than answering occasional chat queries, because no phone battery survives a day of constant cloud round-trips. Qualcomm and MediaTek, the two chipmakers who will actually ship phones on Arm's designs, both face the same pressure from Apple, whose own silicon roadmap has pushed on-device AI performance every generation since the first Neural Engine.

The gap between a platform announcement and a shipped phone

Arm licenses designs; it does not fabricate or sell chips to consumers. Every performance claim in this announcement is Arm's own, unverified by independent silicon until a partner actually builds and ships a device on it, and that gap runs at least a year given Arm's own 2027 timeline. Phone gaming graphics improvements have a mixed track record of translating headline specs into real day-to-day battery life, and "70% speedup" on unspecified small language models is a number without a baseline model size or task attached.

For AI application developers building for mobile, the practical read is to start designing for a split inference architecture now -- lightweight agent steps running locally, heavier reasoning falling back to the cloud -- because that is the architecture every major chipmaker is now building silicon to support, roughly eighteen months before consumers will hold a device that runs it well.

The comparison point nobody in the announcement mentions

Apple has run this exact playbook for four straight generations of its own A-series and M-series silicon, expanding Neural Engine throughput each year specifically so on-device features -- transcription, photo search, and now its own agentic Siri work -- do not depend on a network round-trip. Arm's CSS for Mobile 2 is the rest of the Android ecosystem catching up to that same architectural bet, delivered as a licensable platform rather than owned end-to-end silicon, which is both Arm's advantage -- every non-Apple phone maker can build on it -- and its constraint, since Qualcomm and MediaTek will each tune, brand and ship it on their own timelines rather than Arm's.

ShareXLinkedInEmail

More on

Arm

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.