VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog๐ŸคPartner
Home/Blog/GPT-5.6 Sol on Cerebras: 750 Tokens/Second, OpenAI's $20B Deal, and What It Means for CBRS Stock
AI & TechnologyAugust 5, 2026ยท8 min readยทยทLast updated: 2026-08-12

GPT-5.6 Sol on Cerebras: 750 Tokens/Second, OpenAI's $20B Deal, and What It Means for CBRS Stock

OpenAI's GPT-5.6 Sol runs on Cerebras WSE-3 chips at up to 750 tokens per second, roughly 10x typical Nvidia H100 speed, under a multi-year contract valued above $20 billion.

TC
Trace Cohen
Co-Founder & GP at Six Point Ventures ยท 3x founder (BrandYourself, Launch.it, SPOT) ยท 65+ investments ยท Based in Boca Raton, FL
@Trace_Cohenยทt@nyvp.comยทSouth Florida Advisory
65+Investments3xFounder$200M+Funds Tracked
ShareXLinkedInEmailQuote card

Quick Answer

OpenAI's GPT-5.6 Sol runs on Cerebras WSE-3 chips at up to 750 tokens per second, about 10x faster than typical Nvidia H100 inference, under a contract valued above $20 billion. CBRS stock traded near $238 on August 12, 2026, ahead of Q2 earnings released after market close that day.

OpenAI's GPT-5.6 Sol now runs on Cerebras wafer-scale chips at up to 750 tokens per second โ€” about 10x the roughly 70 tokens per second a typical Nvidia H100 cluster delivers on the same weights. Behind that benchmark is a newly public chipmaker using its biggest customer deal yet to prove its architecture works at the frontier, right as investors wait to see if the revenue shows up.

Cerebras announced on June 26, 2026 that GPT-5.6 Sol โ€” OpenAI's flagship model, reviewed under the White House's voluntary pre-deployment process โ€” would launch on its WSE-3 hardware in July. The deployment is the first time OpenAI has publicly run a frontier model in production on non-Nvidia silicon at this scale, and it lands three months after Cerebras's own IPO, with Q2 earnings due after market close today, August 12.

Close-up of AI accelerator chip hardware representing wafer-scale inference silicon
750 tok/sec
~10x typical H100
Peak Inference Speed
$20B+
multi-year, 750MW
OpenAI Contract Value
~$238
-32% from IPO peak
CBRS Price (Aug 12)
~$194M
+88% YoY
Q2 2026 Revenue Guide

Figures from OpenAI's GPT-5.6 Sol announcement, Cerebras investor guidance issued with Q1 2026 earnings, and stock data as of intraday trading August 12, 2026, ahead of that afternoon's Q2 results (TipRanks, stockanalysis.com).

What Is GPT-5.6 Sol Running on Cerebras?

GPT-5.6 Sol on Cerebras is OpenAI's newest flagship model deployed on Cerebras's WSE-3 wafer-scale processors, delivering up to 750 tokens per second โ€” the fastest production inference speed publicly disclosed for a frontier-class model. Independent estimates put the model at roughly 3 trillion total parameters with 150 billion active per token across 70 layers, reportedly spread one layer per wafer across 70 to 100 Cerebras wafers, according to analysis shared by AI researcher Chubby (@kimmonismus) on X.

That architecture detail matters more than it sounds. Instead of bolting a large model onto whatever inference hardware happens to be available, the layer-per-wafer split suggests OpenAI and Cerebras co-designed the deployment around Cerebras's silicon from early in the model's development, according to OpenAI's own preview announcement. That's a meaningfully deeper commitment than a customer simply renting compute.

How Fast Is 750 Tokens Per Second, Really?

Cerebras's 750 tokens per second on GPT-5.6 Sol compares to roughly 70 tokens per second on a typical Nvidia H100 cluster running the same class of model โ€” Cerebras itself has cited more than 13x GPU throughput on large models in its own benchmarking. Groq, the closest wafer/large-die rival, serves Meta's Llama family at around 500 tokens per second; SambaNova runs older Llama deployments closer to 250 tokens per second. Nvidia GPU clusters generally land in the 40-120 tokens-per-second range for streaming frontier-model completions.

Speed matters most for agentic workloads, which chain many rounds of tool calls and generations end to end. A task that takes 30 seconds on a standard GPU cluster can complete in under 3 seconds at 750 tokens per second โ€” the difference, in practice, between a user waiting out an agent and abandoning it mid-task. That's the pitch Cerebras is making to enterprise buyers evaluating agentic deployments on the AI Valuations Dashboard.

PlatformArchitecturePeak SpeedFrontier Model Win?Public Market Status
Cerebras (WSE-3)Wafer-scale750 tok/sec (GPT-5.6 Sol)Yes โ€” OpenAIPublic (CBRS, Nasdaq)
Groq (LPU)Compiler-scheduled LPU~500 tok/sec (Llama family)NoPrivate
SambaNova (SN40L)Reconfigurable dataflow~250 tok/sec (older Llama)NoPrivate
Nvidia H100 (typical)GPU (Hopper)~70 tok/sec single-streamn/a โ€” incumbentPublic (NVDA)
AWS Trainium2Custom ASICComparable to H100 (AWS claim)NoPublic (AMZN, parent)
AMD MI300XGPU (CDNA 3)Broadly comparable to H100/H200NoPublic (AMD)

Figures blended from Cerebras and OpenAI public disclosures (June-August 2026), Groq and SambaNova publicly reported serving speeds, and AWS's own Trainium2-vs-H100 cost claims. Groq and SambaNova figures were not independently benchmarked by Value Add VC.

What Is Cerebras's OpenAI Contract Actually Worth?

Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute โ€” its largest customer commitment on record, and one it was already marketing to public-market investors ahead of its May 14, 2026 IPO. The GPT-5.6 Sol launch is the first tangible proof point that the contract is converting into a live, high-profile deployment rather than sitting as a backlog figure in an S-1.

Cerebras guided Q2 2026 core revenue to about $194 million, implying 88% year-over-year growth, with full-year 2026 core revenue guided to $855 million to $865 million, according to the company's Q1 2026 investor release. Core gross margin is guided at 36-38%, but core operating margin is still guided negative, at -30% to -32% โ€” a reminder that landing OpenAI as a marquee customer hasn't yet made Cerebras profitable.

What Has CBRS Stock Done Since the IPO?

CBRS closed its May 14, 2026 IPO day up 68% at a roughly $67 billion market cap, spiked as high as $350-plus in the following weeks, then slid to a $160.81 low by June 26 โ€” the same day Cerebras announced the GPT-5.6 Sol deal โ€” before climbing back to trade near $238 intraday on August 12, 2026, a market cap around $68 billion. That puts the stock about 32% below its post-IPO peak but nearly 48% above its June low, heading into Q2 results due after market close that same afternoon.

Does This Change OpenAI's Reliance on Nvidia?

Not dramatically, at least not yet. OpenAI still runs the overwhelming majority of its training and inference workloads on Nvidia GPUs and has separate multi-billion-dollar infrastructure commitments with Nvidia, AMD, and Amazon's Trainium chips. GPT-5.6 Sol on Cerebras is best read as diversification at the margin โ€” a bet that some workloads, especially latency-sensitive agentic ones, are worth routing to specialized wafer-scale hardware even while the bulk of compute stays on Nvidia. Nvidia's own data-center revenue keeps growing well into the tens of billions per quarter, so a single frontier-model deployment on Cerebras doesn't move that needle much in dollar terms.

What it does change is the negotiating leverage in future contracts. Every large AI lab now has at least one credible non-Nvidia option it can point to publicly, and Cerebras just became the most visible proof that a challenger architecture can win a flagship deployment rather than just a benchmark headline.

What the Speed Record Misses

The 750-tokens-per-second headline is real and independently verifiable against Cerebras's prior model deployments, but it's not the whole picture. Initial GPT-5.6 Sol access on Cerebras is limited to a smaller set of customers while capacity scales โ€” this is not yet a general-availability deployment for every OpenAI API user. And the $20 billion contract figure was disclosed as a multi-year commitment, not revenue recognized to date; Cerebras's own Q2 guidance of roughly $194 million in core revenue is a small fraction of that headline number, and its operating margin is still deeply negative. A speed record is a strong marketing proof point ahead of an earnings report, but it doesn't by itself resolve whether Cerebras can convert wafer-scale advantage into sustainable profit at scale โ€” that's what today's numbers, not the tokens-per-second figure, will actually answer.

OpenAI just put its flagship model on non-Nvidia silicon for the first time at production scale.

Whether that's a genuine inference-hardware shift or a one-customer proof point gets tested tonight, when Cerebras reports Q2 earnings after market close.

The Bottom Line on GPT-5.6 Sol and Cerebras

GPT-5.6 Sol running at 750 tokens per second on Cerebras is the clearest evidence yet that wafer-scale inference architecture can win a frontier-model deployment away from Nvidia, and it validates the multi-year, $20 billion-plus OpenAI contract Cerebras marketed ahead of its IPO. CBRS has clawed back to around $238 intraday on August 12 โ€” still about 32% below its post-IPO peak but nearly 48% above its June low โ€” on a company still guiding to a negative 30% operating margin. The real test isn't the tokens-per-second headline: Cerebras reports Q2 2026 results after market close today, and that print will show whether OpenAI-scale demand is actually moving the revenue line.

For now, Cerebras has the fastest disclosed inference speed for a frontier model in production and the most prominent single customer relationship of any AI chip challenger. That's a genuinely strong position three months after an IPO. It's also not yet a profitable one.

Track AI infrastructure valuations and public AI-chip companies on the AI Valuations Dashboard and Tech IPO Tracker at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.

Get VC data most people never see

โ€” 100% free

Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.

ShareXLinkedInEmailQuote card

Frequently Asked Questions

What is GPT-5.6 Sol and why does it run on Cerebras instead of Nvidia?

GPT-5.6 Sol is OpenAI's flagship model, previewed June 26, 2026, and estimated to span roughly 3 trillion total parameters with 150 billion active per token across 70 layers. OpenAI placed it on Cerebras's WSE-3 wafer-scale chips โ€” reportedly one model layer per wafer across 70-100 wafers โ€” because Cerebras's architecture delivers up to 750 tokens per second versus roughly 70 tokens per second on a typical Nvidia H100 cluster running the same weights.

How much is OpenAI's contract with Cerebras worth?

Cerebras has disclosed a multi-year OpenAI contract worth more than $20 billion, covering 750 megawatts of dedicated inference compute. That figure was disclosed before the GPT-5.6 Sol launch and represents Cerebras's largest customer commitment to date, though initial GPT-5.6 Sol access is limited to a smaller set of customers while capacity scales.

Is Cerebras (CBRS) stock a good buy after the OpenAI GPT-5.6 deal?

CBRS traded around $238 on August 12, 2026, giving Cerebras a market cap near $68 billion ahead of that afternoon's Q2 earnings release, on a company still guiding to a core operating margin of roughly -30% for 2026. The stock is down about 32% from its $350-plus post-IPO peak but well above its $160.81 June low, and Cerebras's Q2 2026 results โ€” due after market close August 12 โ€” will show whether the GPT-5.6 Sol deployment is already moving revenue.

How does Cerebras's 750 tokens per second compare to Groq and SambaNova?

Cerebras's 750 tokens per second on GPT-5.6 Sol is faster than Groq's roughly 500 tokens per second serving Meta's Llama models and SambaNova's roughly 250 tokens per second on older Llama deployments. All three are wafer- or large-die inference specialists competing against Nvidia GPU clusters, which typically run frontier models in the 40-120 tokens-per-second range for streaming completions.

Related Tools & Dashboards

๐Ÿค–AI Valuations๐Ÿ“ˆTech IPO Tracker

Keep Reading

๐Ÿ“‰Cerebras Stock Since the IPO: CBRS Down From $386 to Under $170, Then Crashed on Earnings๐ŸŽฏCerebras IPO: What the AI Chip Company's Listing Means for Crossover Investor Confidence๐Ÿ’ฐOpenAI at $300B: How the World's Most Valuable AI Company Is Being Priced

Explore 45+ free VC tools, dashboards, and recommended startup software.

Explore DashboardsHelpful Apps & Platforms

Trace Cohen is a serial founder, investor and data geek. Please feel free to reach out t@nyvp.com

VC
Value Add VC
Helpful AppsSponsor a postTwitterContact