VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog๐ŸคPartner
Home/Blog/GPT-5.6 Sol on Cerebras: 750 Tokens/Second, OpenAI's $20B Deal, and What It Means for CBRS Stock
AI & TechnologyAugust 5, 2026ยท8 min readยท

GPT-5.6 Sol on Cerebras: 750 Tokens/Second, OpenAI's $20B Deal, and What It Means for CBRS Stock

OpenAI's GPT-5.6 Sol runs on Cerebras WSE-3 chips at up to 750 tokens per second, roughly 10x typical Nvidia H100 speed, under a multi-year contract valued above $20 billion.

TC
Trace Cohen
Co-Founder & GP at Six Point Ventures ยท 3x founder (BrandYourself, Launch.it, SPOT) ยท 65+ investments ยท Based in Boca Raton, FL
@Trace_Cohenยทt@nyvp.comยทSouth Florida Advisory
65+Investments3xFounder$200M+Funds Tracked
ShareXLinkedInEmailQuote card

Quick Answer

OpenAI's GPT-5.6 Sol runs on Cerebras WSE-3 chips at up to 750 tokens per second, about 10x faster than typical Nvidia H100 inference, under a contract valued above $20 billion. CBRS stock traded near $224 on August 5, 2026, down 42% from its post-IPO high but up 40% from its June low.

OpenAI's GPT-5.6 Sol now runs on Cerebras wafer-scale chips at up to 750 tokens per second โ€” about 10x the roughly 70 tokens per second a typical Nvidia H100 cluster delivers on the same weights. That's the short answer. The longer answer is a newly public chipmaker using its biggest customer deal yet to prove its architecture works at the frontier, right as investors wait to see if the revenue shows up.

Cerebras announced on June 26, 2026 that GPT-5.6 Sol โ€” OpenAI's flagship model, reviewed under the White House's voluntary pre-deployment process โ€” would launch on its WSE-3 hardware in July. The deployment is the first time OpenAI has publicly run a frontier model in production on non-Nvidia silicon at this scale, and it lands five months after Cerebras's own IPO, with Q2 earnings due August 12.

750 tok/sec
~10x typical H100
Peak Inference Speed
$20B+
multi-year, 750MW
OpenAI Contract Value
~$224.50
-42% from IPO peak
CBRS Price (Aug 5)
~$194M
+88% YoY
Q2 2026 Revenue Guide

Figures from OpenAI's GPT-5.6 Sol announcement, Cerebras investor guidance issued with Q1 2026 earnings, and stock data as of market close August 5, 2026 (stockanalysis.com).

What Is GPT-5.6 Sol Running on Cerebras?

GPT-5.6 Sol on Cerebras is OpenAI's newest flagship model deployed on Cerebras's WSE-3 wafer-scale processors, delivering up to 750 tokens per second โ€” the fastest production inference speed publicly disclosed for a frontier-class model. Independent estimates put the model at roughly 3 trillion total parameters with 150 billion active per token across 70 layers, reportedly spread one layer per wafer across 70 to 100 Cerebras wafers, according to analysis shared by AI researcher Chubby (@kimmonismus) on X.

That architecture detail matters more than it sounds. Instead of bolting a large model onto whatever inference hardware happens to be available, the layer-per-wafer split suggests OpenAI and Cerebras co-designed the deployment around Cerebras's silicon from early in the model's development, according to OpenAI's own preview announcement. That's a meaningfully deeper commitment than a customer simply renting compute.

How Fast Is 750 Tokens Per Second, Really?

Cerebras's 750 tokens per second on GPT-5.6 Sol compares to roughly 70 tokens per second on a typical Nvidia H100 cluster running the same class of model โ€” Cerebras itself has cited more than 13x GPU throughput on large models in its own benchmarking. Groq, the closest wafer/large-die rival, serves Meta's Llama family at around 500 tokens per second; SambaNova runs older Llama deployments closer to 250 tokens per second. Nvidia GPU clusters generally land in the 40-120 tokens-per-second range for streaming frontier-model completions.

Speed matters most for agentic workloads, which chain many rounds of tool calls and generations end to end. A task that takes 30 seconds on a standard GPU cluster can complete in under 3 seconds at 750 tokens per second โ€” the difference, in practice, between a user waiting out an agent and abandoning it mid-task. That's the pitch Cerebras is making to enterprise buyers evaluating agentic deployments on the AI Valuations Dashboard.

What Is Cerebras's OpenAI Contract Actually Worth?

Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute โ€” its largest customer commitment on record, and one it was already marketing to public-market investors ahead of its May 14, 2026 IPO. The GPT-5.6 Sol launch is the first tangible proof point that the contract is converting into a live, high-profile deployment rather than sitting as a backlog figure in an S-1.

Cerebras guided Q2 2026 core revenue to about $194 million, implying 88% year-over-year growth, with full-year 2026 core revenue guided to $855 million to $865 million, according to the company's Q1 2026 investor release. Core gross margin is guided at 36-38%, but core operating margin is still guided negative, at -30% to -32% โ€” a reminder that landing OpenAI as a marquee customer hasn't yet made Cerebras profitable.

What Has CBRS Stock Done Since the IPO?

CBRS closed its May 14, 2026 IPO day up 68% at a roughly $67 billion market cap, then slid to a $160.81 low by June 26 โ€” the same day Cerebras announced the GPT-5.6 Sol deal โ€” before recovering to trade near $224.50 on August 5, 2026, a market cap around $69.7 billion. That puts the stock about 42% below its $386.34 post-IPO peak but roughly 40% above its June low, with a trailing P/E above 190x reflecting how much future growth is already priced in.

Does This Change OpenAI's Reliance on Nvidia?

Not dramatically, at least not yet. OpenAI still runs the overwhelming majority of its training and inference workloads on Nvidia GPUs and has separate multi-billion-dollar infrastructure commitments with Nvidia, AMD, and Amazon's Trainium chips. GPT-5.6 Sol on Cerebras is best read as diversification at the margin โ€” a bet that some workloads, especially latency-sensitive agentic ones, are worth routing to specialized wafer-scale hardware even while the bulk of compute stays on Nvidia. Nvidia's own data-center revenue keeps growing well into the tens of billions per quarter, so a single frontier-model deployment on Cerebras doesn't move that needle much in dollar terms.

What it does change is the negotiating leverage in future contracts. Every large AI lab now has at least one credible non-Nvidia option it can point to publicly, and Cerebras just became the most visible proof that a challenger architecture can win a flagship deployment rather than just a benchmark headline.

What the Speed Record Misses

The 750-tokens-per-second headline is real and independently verifiable against Cerebras's prior model deployments, but it's not the whole picture. Initial GPT-5.6 Sol access on Cerebras is limited to a smaller set of customers while capacity scales โ€” this is not yet a general-availability deployment for every OpenAI API user. And the $20 billion contract figure was disclosed as a multi-year commitment, not revenue recognized to date; Cerebras's own Q2 guidance of roughly $194 million in core revenue is a small fraction of that headline number, and its operating margin is still deeply negative. A speed record is a strong marketing proof point ahead of an August 12 earnings report, but it doesn't by itself resolve whether Cerebras can convert wafer-scale advantage into sustainable profit at scale.

OpenAI just put its flagship model on non-Nvidia silicon for the first time at production scale.

Whether that's a genuine inference-hardware shift or a one-customer proof point gets tested on August 12.

The Bottom Line on GPT-5.6 Sol and Cerebras

GPT-5.6 Sol running at 750 tokens per second on Cerebras is the clearest evidence yet that wafer-scale inference architecture can win a frontier-model deployment away from Nvidia, and it validates the multi-year, $20 billion-plus OpenAI contract Cerebras marketed ahead of its IPO. But CBRS still trades at a triple-digit P/E on a company guiding to a negative 30% operating margin, and the real test isn't the tokens-per-second headline โ€” it's whether Cerebras's August 12 earnings show that OpenAI-scale demand is actually moving the revenue line.

For now, Cerebras has the fastest disclosed inference speed for a frontier model in production and the most prominent single customer relationship of any AI chip challenger. That's a genuinely strong position three months after an IPO. It's also not yet a profitable one.

Track AI infrastructure valuations and public AI-chip companies on the AI Valuations Dashboard and Tech IPO Tracker at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.

Get VC data most people never see

โ€” 100% free

Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.

ShareXLinkedInEmailQuote card

Frequently Asked Questions

What is GPT-5.6 Sol and why does it run on Cerebras instead of Nvidia?

GPT-5.6 Sol is OpenAI's flagship model, previewed June 26, 2026, and estimated to span roughly 3 trillion total parameters with 150 billion active per token across 70 layers. OpenAI placed it on Cerebras's WSE-3 wafer-scale chips โ€” reportedly one model layer per wafer across 70-100 wafers โ€” because Cerebras's architecture delivers up to 750 tokens per second versus roughly 70 tokens per second on a typical Nvidia H100 cluster running the same weights.

How much is OpenAI's contract with Cerebras worth?

Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute. That figure was disclosed before the GPT-5.6 Sol launch and represents Cerebras's largest customer commitment to date, though initial GPT-5.6 Sol access is limited to a smaller set of customers while capacity scales.

Is Cerebras (CBRS) stock a good buy after the OpenAI GPT-5.6 deal?

CBRS traded around $224.50 on August 5, 2026, giving Cerebras a market cap near $69.7 billion and a trailing P/E above 190x โ€” a steep valuation for a company still guiding to a core operating margin of roughly -30% for 2026. The stock is down about 42% from its $386 post-IPO peak but up roughly 40% from its $160.81 June low, and Cerebras reports Q2 2026 earnings on August 12, which will show whether the GPT-5.6 Sol deployment is already moving revenue.

How does Cerebras's 750 tokens per second compare to Groq and SambaNova?

Cerebras's 750 tokens per second on GPT-5.6 Sol is faster than Groq's roughly 500 tokens per second serving Meta's Llama models and SambaNova's roughly 250 tokens per second on older Llama deployments. All three are wafer- or large-die inference specialists competing against Nvidia GPU clusters, which typically run frontier models in the 40-120 tokens-per-second range for streaming completions.

Related Tools & Dashboards

๐Ÿค–AI Valuations๐Ÿ“ˆTech IPO Tracker

Keep Reading

๐Ÿ“‰Cerebras Stock Since the IPO: CBRS Down From $386 to Under $170, Then Crashed on Earnings๐ŸŽฏCerebras IPO: What the AI Chip Company's Listing Means for Crossover Investor Confidence๐Ÿ’ฐOpenAI at $300B: How the World's Most Valuable AI Company Is Being Priced

Explore 45+ free VC tools, dashboards, and recommended startup software.

Explore DashboardsHelpful Apps & Platforms

Trace Cohen is a serial founder, investor and data geek. Please feel free to reach out t@nyvp.com

VC
Value Add VC
Helpful AppsSponsor a postTwitterContact