OpenAI's GPT-5.6 Sol now runs on Cerebras wafer-scale chips at up to 750 tokens per second โ about 10x the roughly 70 tokens per second a typical Nvidia H100 cluster delivers on the same weights. That's the short answer. The longer answer is a newly public chipmaker using its biggest customer deal yet to prove its architecture works at the frontier, right as investors wait to see if the revenue shows up.
Cerebras announced on June 26, 2026 that GPT-5.6 Sol โ OpenAI's flagship model, reviewed under the White House's voluntary pre-deployment process โ would launch on its WSE-3 hardware in July. The deployment is the first time OpenAI has publicly run a frontier model in production on non-Nvidia silicon at this scale, and it lands five months after Cerebras's own IPO, with Q2 earnings due August 12.
Figures from OpenAI's GPT-5.6 Sol announcement, Cerebras investor guidance issued with Q1 2026 earnings, and stock data as of market close August 5, 2026 (stockanalysis.com).
What Is GPT-5.6 Sol Running on Cerebras?
GPT-5.6 Sol on Cerebras is OpenAI's newest flagship model deployed on Cerebras's WSE-3 wafer-scale processors, delivering up to 750 tokens per second โ the fastest production inference speed publicly disclosed for a frontier-class model. Independent estimates put the model at roughly 3 trillion total parameters with 150 billion active per token across 70 layers, reportedly spread one layer per wafer across 70 to 100 Cerebras wafers, according to analysis shared by AI researcher Chubby (@kimmonismus) on X.
That architecture detail matters more than it sounds. Instead of bolting a large model onto whatever inference hardware happens to be available, the layer-per-wafer split suggests OpenAI and Cerebras co-designed the deployment around Cerebras's silicon from early in the model's development, according to OpenAI's own preview announcement. That's a meaningfully deeper commitment than a customer simply renting compute.
How Fast Is 750 Tokens Per Second, Really?
Cerebras's 750 tokens per second on GPT-5.6 Sol compares to roughly 70 tokens per second on a typical Nvidia H100 cluster running the same class of model โ Cerebras itself has cited more than 13x GPU throughput on large models in its own benchmarking. Groq, the closest wafer/large-die rival, serves Meta's Llama family at around 500 tokens per second; SambaNova runs older Llama deployments closer to 250 tokens per second. Nvidia GPU clusters generally land in the 40-120 tokens-per-second range for streaming frontier-model completions.
Speed matters most for agentic workloads, which chain many rounds of tool calls and generations end to end. A task that takes 30 seconds on a standard GPU cluster can complete in under 3 seconds at 750 tokens per second โ the difference, in practice, between a user waiting out an agent and abandoning it mid-task. That's the pitch Cerebras is making to enterprise buyers evaluating agentic deployments on the AI Valuations Dashboard.
What Is Cerebras's OpenAI Contract Actually Worth?
Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute โ its largest customer commitment on record, and one it was already marketing to public-market investors ahead of its May 14, 2026 IPO. The GPT-5.6 Sol launch is the first tangible proof point that the contract is converting into a live, high-profile deployment rather than sitting as a backlog figure in an S-1.
Cerebras guided Q2 2026 core revenue to about $194 million, implying 88% year-over-year growth, with full-year 2026 core revenue guided to $855 million to $865 million, according to the company's Q1 2026 investor release. Core gross margin is guided at 36-38%, but core operating margin is still guided negative, at -30% to -32% โ a reminder that landing OpenAI as a marquee customer hasn't yet made Cerebras profitable.
What Has CBRS Stock Done Since the IPO?
CBRS closed its May 14, 2026 IPO day up 68% at a roughly $67 billion market cap, then slid to a $160.81 low by June 26 โ the same day Cerebras announced the GPT-5.6 Sol deal โ before recovering to trade near $224.50 on August 5, 2026, a market cap around $69.7 billion. That puts the stock about 42% below its $386.34 post-IPO peak but roughly 40% above its June low, with a trailing P/E above 190x reflecting how much future growth is already priced in.
Does This Change OpenAI's Reliance on Nvidia?
Not dramatically, at least not yet. OpenAI still runs the overwhelming majority of its training and inference workloads on Nvidia GPUs and has separate multi-billion-dollar infrastructure commitments with Nvidia, AMD, and Amazon's Trainium chips. GPT-5.6 Sol on Cerebras is best read as diversification at the margin โ a bet that some workloads, especially latency-sensitive agentic ones, are worth routing to specialized wafer-scale hardware even while the bulk of compute stays on Nvidia. Nvidia's own data-center revenue keeps growing well into the tens of billions per quarter, so a single frontier-model deployment on Cerebras doesn't move that needle much in dollar terms.
What it does change is the negotiating leverage in future contracts. Every large AI lab now has at least one credible non-Nvidia option it can point to publicly, and Cerebras just became the most visible proof that a challenger architecture can win a flagship deployment rather than just a benchmark headline.
What the Speed Record Misses
The 750-tokens-per-second headline is real and independently verifiable against Cerebras's prior model deployments, but it's not the whole picture. Initial GPT-5.6 Sol access on Cerebras is limited to a smaller set of customers while capacity scales โ this is not yet a general-availability deployment for every OpenAI API user. And the $20 billion contract figure was disclosed as a multi-year commitment, not revenue recognized to date; Cerebras's own Q2 guidance of roughly $194 million in core revenue is a small fraction of that headline number, and its operating margin is still deeply negative. A speed record is a strong marketing proof point ahead of an August 12 earnings report, but it doesn't by itself resolve whether Cerebras can convert wafer-scale advantage into sustainable profit at scale.
OpenAI just put its flagship model on non-Nvidia silicon for the first time at production scale.
Whether that's a genuine inference-hardware shift or a one-customer proof point gets tested on August 12.
The Bottom Line on GPT-5.6 Sol and Cerebras
GPT-5.6 Sol running at 750 tokens per second on Cerebras is the clearest evidence yet that wafer-scale inference architecture can win a frontier-model deployment away from Nvidia, and it validates the multi-year, $20 billion-plus OpenAI contract Cerebras marketed ahead of its IPO. But CBRS still trades at a triple-digit P/E on a company guiding to a negative 30% operating margin, and the real test isn't the tokens-per-second headline โ it's whether Cerebras's August 12 earnings show that OpenAI-scale demand is actually moving the revenue line.
For now, Cerebras has the fastest disclosed inference speed for a frontier model in production and the most prominent single customer relationship of any AI chip challenger. That's a genuinely strong position three months after an IPO. It's also not yet a profitable one.
Track AI infrastructure valuations and public AI-chip companies on the AI Valuations Dashboard and Tech IPO Tracker at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.
Get VC data most people never see
โ 100% free
Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.