Illustration for: GLM-5.3 Hits the API at $1.40/$4.40 Per Million Tokens

GLM-5.3 Hits the API at $1.40/$4.40 Per Million Tokens

Z.ai priced its new GLM-5.3 model identically to GLM-5.2 at $1.40 per million input tokens and $4.40 per million output tokens, undercutting Grok, Kimi, Claude Opus and GPT on a simple blended-cost comparison.

By the Numbers

$1.40/M tokens
GLM-5.3 input price
$4.40/M tokens
GLM-5.3 output price
$5.80
Blended cost (1M+1M)
$8.00
Grok 4.6 blended cost
$18.00
Kimi K3 blended cost
TC
Early-stage VC & angel · Founder, New York Venture Partners · Value Add Pulse AI Desk
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

GLM-5.3 is priced identically to its predecessor GLM-5.2 -- Z.ai held pricing flat despite shipping a materially upgraded model, an unusual choice in a market where most labs raise prices with each frontier release

2

On a simple blended cost of one million input plus one million output tokens, GLM-5.3 comes to $5.80, versus $8 for Grok 4.6, $18 for Kimi K3, $30 for Claude Opus 5 and $35 for GPT-5.6 Sol

3

GLM-5.3 is reportedly more verbose than its predecessor, meaning flat per-token pricing doesn't guarantee flat costs for a completed task -- a longer average response can offset a chunk of the headline savings

4

The model also shipped with new cyber-capability testing, reportedly identifying a vulnerability in a widely used coding tool during evaluation

TC

The VC Read · Trace's Take

Trace Cohen

Flat pricing on a capability upgrade is a share-grab move, not a cost-of-goods story -- Z.ai is betting volume growth from undercutting Western labs matters more than near-term margin. Anyone benchmarking model costs for a production workload should run their own actual prompts through both models before switching on headline pricing alone, because the verbosity gap can erase a meaningful chunk of the $2.20 blended-cost advantage on paper.

Analysis

Z.ai's newest model, GLM-5.3, is now available through its API at $1.40 per million input tokens and $4.40 per million output tokens -- exactly the same rate as its predecessor, GLM-5.2, VentureBeat reported. Holding pricing flat on a materially upgraded model release is a departure from the pattern most frontier labs have followed in 2026, where price increases have generally tracked capability jumps.

On VentureBeat's simple blended-cost comparison -- one million input tokens plus one million output tokens -- GLM-5.3 works out to $5.80, compared with $8 for Grok 4.6 at its lower context rate, $18 for Kimi K3, $30 for Claude Opus 5, and $35 for GPT-5.6 Sol. That places GLM-5.3 meaningfully below every major Western and Chinese frontier competitor on this specific comparison, continuing a trend of Chinese open-weight and hosted models undercutting US frontier-lab pricing by a wide margin.

“Buyers comparing models purely on headline per-token price without normalizing for output length risk underestimating actual production costs.”

The comparison comes with an important caveat: GLM-5.3 is reportedly more verbose than its predecessor, meaning it tends to generate longer responses for a comparable task. Flat per-token pricing on a more verbose model doesn't necessarily translate into flat real-world costs for a completed workload, since a longer average output consumes more of the (more expensive) output-token budget even at an unchanged per-token rate. Buyers comparing models purely on headline per-token price without normalizing for output length risk underestimating actual production costs.

GLM-5.3 also shipped with expanded cyber-capability testing; according to reporting, evaluators found the model capable of identifying a serious vulnerability in Cursor during testing. Chinese open-weight labs are now closing the gap with Western frontier models on specialized technical benchmarks, not only on price.

For founders building AI-native products, the pricing gap between GLM-5.3 and Western frontier models is now wide enough to change default vendor choices for cost-sensitive, high-volume workloads like customer support automation or bulk document processing, even if teams still reach for Claude or GPT for tasks where reasoning quality matters more than throughput cost. Z.ai has not disclosed usage or revenue figures for GLM-5.3, so it's not yet clear how much of this pricing advantage is translating into actual developer adoption outside of China.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take, a few times a week. Free to subscribe, no spam.

Read Next

AI

Z.ai Launches ZCode, a Free IDE for GLM-5.2, to Challenge Cursor, Claude Code and Copilot

Illustration for: Z.ai Launches ZCode, a Free IDE for GLM-5.2, to Challenge Cursor, Claude Code and Copilot
AIUp to 82% cheaper than Opus 4.8

Z.ai Launches ZCode, a Free IDE for GLM-5.2, to Challenge Cursor, Claude Code and Copilot

Z.ai, the Beijing lab formerly known as Zhipu AI, officially launched ZCode on July 2 — a free, agent-first development environment purpose-built for its GLM-5.2 model, with subscription plans starting at $16.20/month, up to 82% cheaper than Anthropic's Claude Opus 4.8 API pricing. The launch lands three weeks after a US export-control order briefly suspended foreign access to Anthropic's Claude Fable 5 and Mythos 5, an episode Z.ai's timing has directly capitalized on, with GLM-5.2 running entirely on Huawei silicon and ranking second globally on the Code Arena leaderboard behind only Claude Fable 5.

AI

Chinese Lab Z.ai Launches ZCode to Challenge Cursor, Claude Code and GitHub Copilot

Illustration for: Chinese Lab Z.ai Launches ZCode to Challenge Cursor, Claude Code and GitHub Copilot
AINew product launch

Chinese Lab Z.ai Launches ZCode to Challenge Cursor, Claude Code and GitHub Copilot

Z.ai launched ZCode on July 2, an agentic development environment built around its GLM Coding Plan and long-horizon task execution — the user describes an outcome, the agent plans the work, edits files, runs checks and iterates until the goal is met. It enters a coding-agent market where, per VentureBeat, virtually every developer shipping software in 2026 is already using Claude Code, GitHub Copilot or Cursor, positioning ZCode as a lower-cost, desktop-native challenger.

IPO

China's AI Labs Race Toward A Wave Of IPOs

Illustration for: China's AI Labs Race Toward A Wave Of IPOs
IPO

China's AI Labs Race Toward A Wave Of IPOs

Moonshot, DeepSeek, MiniMax and Z.ai are all pursuing or completing public listings within months of each other, a compressed wave of Chinese AI IPOs that resets how investors value the entire sector.