VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog
โ† Value Add PulseAI$15/1M output tokens

Moonshot's Kimi K3 Claims Near-Frontier Performance

Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks third on Artificial Analysis's GDPval-AA v2 benchmark behind only Claude Fable 5 Max and GPT-5.6 Sol Max, with open weights promised by July 27.

2.8T total
Parameters
1.05M tokens
Context window
3rd overall
GDPval-AA v2 rank
$15 / 1M tokens
Output pricing
By July 27
Open-weight release
TC
Trace Cohen
Early-stage VC & angel ยท Founder, New York Venture Partners
July 16, 2026
1 min read
ShareXLinkedInEmail
THE RUNDOWN
1

On GDPval-AA v2, a benchmark measuring real-world tasks across 44 occupations, Kimi K3 scored 1,687 -- third overall behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600)

2

Kimi K3 topped Arena.AI's Frontend Code Arena outright with a score of 1,679, outpacing both Claude Fable 5 and GPT-5.6 Sol on that specific benchmark

3

Moonshot claims roughly 2.5x improvement in scaling efficiency over Kimi K2, via two architectural changes -- Kimi Delta Attention (hybrid linear attention) and Attention Residuals -- rather than simply adding parameters

4

API pricing undercuts frontier U.S. models significantly: $0.30 per million cache-hit input tokens, $3 per million on cache misses, $15 per million output tokens; open weights are promised by July 27

TC
The VC Read ยท Trace's TakeTrace Cohen

The benchmark scores matter less than the pricing -- $15 per million output tokens for third-best-in-class performance rewrites the unit economics every AI-native startup has been modeling around GPT-5.6 or Claude pricing. If Kimi K3's open weights ship as promised on July 27, expect a wave of infra startups to quietly add it as a fallback or default model within weeks, the same way DeepSeek got absorbed into serving stacks last year. The labs that should be most worried aren't the frontier three, it's every mid-tier model provider whose whole pitch was 'cheaper than GPT, almost as good' -- Kimi K3 just did that better.

Moonshot AI unveiled Kimi K3 at Shanghai's World AI Conference, a 2.8-trillion-parameter model with a 1.05-million-token context window that the company says is its most capable to date. On Artificial Analysis's GDPval-AA v2 benchmark -- which measures real-world task performance across 44 occupations and nine major industries -- Kimi K3 scored 1,687, placing third overall behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600). On AA-Briefcase, a private agentic benchmark testing long-horizon knowledge work, K3 climbed to second place at 1,527, beating GPT-5.6 Sol Max's 1,495. It took the No. 1 spot outright on Arena.AI's Frontend Code Arena.

Moonshot attributes roughly a 2.5x improvement in scaling efficiency over its prior Kimi K2 model to two architectural changes: Kimi Delta Attention, a hybrid linear-attention scheme, and Attention Residuals, which change how information moves between transformer layers -- efficiency gains rather than pure parameter-count scaling, echoing the efficiency-first approach that made DeepSeek's earlier releases notable.

โ€œAI infrastructure spending has outrun what's actually needed to stay competitive.โ€

Pricing is aggressive relative to frontier U.S. models: $0.30 per million cache-hit input tokens, $3 per million on cache misses, and $15 per million output tokens, available now via API with open weights promised by July 27. That combination of near-frontier benchmark performance and a fraction of the price is precisely what triggered the broader chip-stock selloff this week, on fears it signals U.S. AI infrastructure spending has outrun what's actually needed to stay competitive.

For developers and enterprises, Kimi K3's open-weight release later this month will matter more than the benchmark scores alone -- it gives any company the ability to self-host a near-frontier model rather than depend on API access from a small number of U.S. labs. What to watch next: independent benchmark verification once the open weights ship, and whether Western enterprises adopt K3 given ongoing data-sovereignty and export-control considerations around Chinese-origin models.

ShareXLinkedInEmail
More onAnthropic โ†’

Originally reported by VentureBeat. Analysis and editorial commentary by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI$10B compute lease

Anthropic in Talks to Lease $10B of Meta's Compute

Anthropic and Meta are in early talks for Anthropic to lease up to $10B of Meta's AI compute over two years, letting Meta monetize its buildout while Anthropic diversifies beyond Amazon and Google.

AI

China's Open-Weight Wave Forces an Enterprise Rethink

Kimi K3's benchmark-topping debut is accelerating enterprise interest in open-weight models, forcing US buyers to weigh self-hosted Chinese models against closed, subscription-priced offerings from Anthropic and OpenAI.

AI$960,000

Jensen Huang's Leather Jacket Sells for $960K

A leather jacket worn by Nvidia CEO Jensen Huang sold for $960,000 at Sotheby's, nearly 20 times its pre-sale estimate, with proceeds benefiting a philanthropic initiative for young tech builders.

@Trace_Cohenยทt@nyvp.com