OpenAI's GPT-5.6 Sol now runs on Cerebras wafer-scale chips at up to 750 tokens per second โ about 10x the roughly 70 tokens per second a typical Nvidia H100 cluster delivers on the same weights. Behind that benchmark is a newly public chipmaker using its biggest customer deal yet to prove its architecture works at the frontier, right as investors wait to see if the revenue shows up.
Cerebras announced on June 26, 2026 that GPT-5.6 Sol โ OpenAI's flagship model, reviewed under the White House's voluntary pre-deployment process โ would launch on its WSE-3 hardware in July. The deployment is the first time OpenAI has publicly run a frontier model in production on non-Nvidia silicon at this scale, and it lands three months after Cerebras's own IPO, with Q2 earnings due after market close today, August 12.

Figures from OpenAI's GPT-5.6 Sol announcement, Cerebras investor guidance issued with Q1 2026 earnings, and stock data as of intraday trading August 12, 2026, ahead of that afternoon's Q2 results (TipRanks, stockanalysis.com).
What Is GPT-5.6 Sol Running on Cerebras?
GPT-5.6 Sol on Cerebras is OpenAI's newest flagship model deployed on Cerebras's WSE-3 wafer-scale processors, delivering up to 750 tokens per second โ the fastest production inference speed publicly disclosed for a frontier-class model. Independent estimates put the model at roughly 3 trillion total parameters with 150 billion active per token across 70 layers, reportedly spread one layer per wafer across 70 to 100 Cerebras wafers, according to analysis shared by AI researcher Chubby (@kimmonismus) on X.
That architecture detail matters more than it sounds. Instead of bolting a large model onto whatever inference hardware happens to be available, the layer-per-wafer split suggests OpenAI and Cerebras co-designed the deployment around Cerebras's silicon from early in the model's development, according to OpenAI's own preview announcement. That's a meaningfully deeper commitment than a customer simply renting compute.
How Fast Is 750 Tokens Per Second, Really?
Cerebras's 750 tokens per second on GPT-5.6 Sol compares to roughly 70 tokens per second on a typical Nvidia H100 cluster running the same class of model โ Cerebras itself has cited more than 13x GPU throughput on large models in its own benchmarking. Groq, the closest wafer/large-die rival, serves Meta's Llama family at around 500 tokens per second; SambaNova runs older Llama deployments closer to 250 tokens per second. Nvidia GPU clusters generally land in the 40-120 tokens-per-second range for streaming frontier-model completions.
Speed matters most for agentic workloads, which chain many rounds of tool calls and generations end to end. A task that takes 30 seconds on a standard GPU cluster can complete in under 3 seconds at 750 tokens per second โ the difference, in practice, between a user waiting out an agent and abandoning it mid-task. That's the pitch Cerebras is making to enterprise buyers evaluating agentic deployments on the AI Valuations Dashboard.
| Platform | Architecture | Peak Speed | Frontier Model Win? | Public Market Status |
|---|---|---|---|---|
| Cerebras (WSE-3) | Wafer-scale | 750 tok/sec (GPT-5.6 Sol) | Yes โ OpenAI | Public (CBRS, Nasdaq) |
| Groq (LPU) | Compiler-scheduled LPU | ~500 tok/sec (Llama family) | No | Private |
| SambaNova (SN40L) | Reconfigurable dataflow | ~250 tok/sec (older Llama) | No | Private |
| Nvidia H100 (typical) | GPU (Hopper) | ~70 tok/sec single-stream | n/a โ incumbent | Public (NVDA) |
| AWS Trainium2 | Custom ASIC | Comparable to H100 (AWS claim) | No | Public (AMZN, parent) |
| AMD MI300X | GPU (CDNA 3) | Broadly comparable to H100/H200 | No | Public (AMD) |
Figures blended from Cerebras and OpenAI public disclosures (June-August 2026), Groq and SambaNova publicly reported serving speeds, and AWS's own Trainium2-vs-H100 cost claims. Groq and SambaNova figures were not independently benchmarked by Value Add VC.
What Is Cerebras's OpenAI Contract Actually Worth?
Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute โ its largest customer commitment on record, and one it was already marketing to public-market investors ahead of its May 14, 2026 IPO. The GPT-5.6 Sol launch is the first tangible proof point that the contract is converting into a live, high-profile deployment rather than sitting as a backlog figure in an S-1.
Cerebras guided Q2 2026 core revenue to about $194 million, implying 88% year-over-year growth, with full-year 2026 core revenue guided to $855 million to $865 million, according to the company's Q1 2026 investor release. Core gross margin is guided at 36-38%, but core operating margin is still guided negative, at -30% to -32% โ a reminder that landing OpenAI as a marquee customer hasn't yet made Cerebras profitable.
What Has CBRS Stock Done Since the IPO?
CBRS closed its May 14, 2026 IPO day up 68% at a roughly $67 billion market cap, spiked as high as $350-plus in the following weeks, then slid to a $160.81 low by June 26 โ the same day Cerebras announced the GPT-5.6 Sol deal โ before climbing back to trade near $238 intraday on August 12, 2026, a market cap around $68 billion. That puts the stock about 32% below its post-IPO peak but nearly 48% above its June low, heading into Q2 results due after market close that same afternoon.
Does This Change OpenAI's Reliance on Nvidia?
Not dramatically, at least not yet. OpenAI still runs the overwhelming majority of its training and inference workloads on Nvidia GPUs and has separate multi-billion-dollar infrastructure commitments with Nvidia, AMD, and Amazon's Trainium chips. GPT-5.6 Sol on Cerebras is best read as diversification at the margin โ a bet that some workloads, especially latency-sensitive agentic ones, are worth routing to specialized wafer-scale hardware even while the bulk of compute stays on Nvidia. Nvidia's own data-center revenue keeps growing well into the tens of billions per quarter, so a single frontier-model deployment on Cerebras doesn't move that needle much in dollar terms.
What it does change is the negotiating leverage in future contracts. Every large AI lab now has at least one credible non-Nvidia option it can point to publicly, and Cerebras just became the most visible proof that a challenger architecture can win a flagship deployment rather than just a benchmark headline.
What the Speed Record Misses
The 750-tokens-per-second headline is real and independently verifiable against Cerebras's prior model deployments, but it's not the whole picture. Initial GPT-5.6 Sol access on Cerebras is limited to a smaller set of customers while capacity scales โ this is not yet a general-availability deployment for every OpenAI API user. And the $20 billion contract figure was disclosed as a multi-year commitment, not revenue recognized to date; Cerebras's own Q2 guidance of roughly $194 million in core revenue is a small fraction of that headline number, and its operating margin is still deeply negative. A speed record is a strong marketing proof point ahead of an earnings report, but it doesn't by itself resolve whether Cerebras can convert wafer-scale advantage into sustainable profit at scale โ that's what today's numbers, not the tokens-per-second figure, will actually answer.
OpenAI just put its flagship model on non-Nvidia silicon for the first time at production scale.
Whether that's a genuine inference-hardware shift or a one-customer proof point gets tested tonight, when Cerebras reports Q2 earnings after market close.
The Bottom Line on GPT-5.6 Sol and Cerebras
GPT-5.6 Sol running at 750 tokens per second on Cerebras is the clearest evidence yet that wafer-scale inference architecture can win a frontier-model deployment away from Nvidia, and it validates the multi-year, $20 billion-plus OpenAI contract Cerebras marketed ahead of its IPO. CBRS has clawed back to around $238 intraday on August 12 โ still about 32% below its post-IPO peak but nearly 48% above its June low โ on a company still guiding to a negative 30% operating margin. The real test isn't the tokens-per-second headline: Cerebras reports Q2 2026 results after market close today, and that print will show whether OpenAI-scale demand is actually moving the revenue line.
For now, Cerebras has the fastest disclosed inference speed for a frontier model in production and the most prominent single customer relationship of any AI chip challenger. That's a genuinely strong position three months after an IPO. It's also not yet a profitable one.
Track AI infrastructure valuations and public AI-chip companies on the AI Valuations Dashboard and Tech IPO Tracker at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.
Get VC data most people never see
โ 100% free
Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.