AI & TechnologyJuly 12, 2026·13 min read··Last updated: 2026-09-29

Best AI Models (September 2026): Opus 5.5 vs GPT-6, Ranked

Claude Opus 5.5 takes #1 on LMArena and on independent Terminal-Bench 4.0, Claude Fable 5.1 and GPT-6 Astra follow, and Gemini 3.8 Flash is the cheapest model in the top 10.

VC
Editor-in-chief: Trace Cohen — Angel investor, VC, family office, operator and founder · 3x founder (BrandYourself, Launch.it, SPOT) · 65+ investments

AI-assisted: drafted with AI from the cited sources — how we check it

65+Investments3xFounder$200M+Funds Tracked

Quick Answer

As of September 26, 2026, Claude Opus 5.5 is the best AI model overall: #1 on LMArena's text leaderboard at 1509 and #1 on Vals AI's independent Terminal-Bench 4.0 run at 61.6% (53.5% if tasks served by fallback Claude models count as failures), for $4/$20 per million tokens. Claude Fable 5.1 and GPT-6 Astra ($10/$50 each) follow, and Gemini 3.8 Flash is the cheapest model in LMArena's top 10.

Claude Opus 5.5 is the best AI model as of September 26, 2026: #1 on LMArena at 1509 and #1 on an independent Terminal-Bench 4.0 run at 61.6%, at $4/$20 per million tokens. That is less than half the list price of Claude Fable 5.1 or GPT-6 Astra, the two models just behind it. Nine models make this list, and picking the wrong one for the job, not the wrong vendor, is where most teams now lose money.

I evaluate AI tools across 65+ portfolio companies, and September 2026 was the busiest month for frontier releases this year. In four weeks Anthropic shipped Claude Fable 5.1 and Claude Opus 5.5, OpenAI shipped GPT-6 Astra followed by GPT-6 Sol and Luna, Google shipped Gemini 3.8 Flash, Meta shipped Muse Spark 1.3, and xAI shipped Grok 4.7. Four of the seven models in this ranking's August 31 edition have been superseded.

Abstract visualization of neural network nodes representing frontier AI model comparison
9
from Anthropic, OpenAI, Google, Meta and xAI
Models ranked
1509
Claude Opus 5.5, board data Sep 25, 2026
Top LMArena text score
61.6%
Claude Opus 5.5 (53.5% excluding fallback)
Top Terminal-Bench 4.0 (Vals AI)
$3.75–$50
Gemini 3.8 Flash (intro) to Fable 5.1 and GPT-6 Astra
Price range (output/1M tokens)

What Are the Best AI Models in September 2026, Ranked?

Claude Opus 5.5 from Anthropic is the top-ranked AI model as of September 26, 2026. It leads the LMArena text leaderboard at 1509 (board data dated September 25, 2026, read September 26) and the Vals AI Terminal-Bench 4.0 leaderboard at 61.62% (updated September 22, 2026). Claude Fable 5.1 ranks second and GPT-6 Astra third. Gemini 3.8 Flash is the cheapest model in LMArena's top 10.

Meta's Muse Spark 1.3, Google's Gemini 3.1 Pro, Claude Sonnet 5, OpenAI's GPT-6 Sol and xAI's Grok 4.7 round out the nine. The top of LMArena is tight: Opus 5.5's 1509 carries a ±12 confidence interval, and the superseded Claude Fable 5 still sits #3 at 1504. Treat a gap of a few points as a tie. Track the category leader for your specific task, not the vendor brand, when picking a model for a new build.

The 9 Best AI Models in 2026, Ranked (as of September 26)

1
Claude Opus 5.5 (Anthropic)
Released September 22, 2026 at $4 input / $20 output per million tokens, 20% below Opus 5, with a 1M-token context window. #1 on LMArena text at 1509. Scores 61.62% on Vals AI's independent Terminal-Bench 4.0 run, or 53.54% if the 30 of 198 task attempts served by fallback Claude models are counted as failures (Anthropic reports 66.4%), plus 81.8% on OSWorld 2.0 (partial) and 67.7% on Humanity's Last Exam with tools. Anthropic says it 'performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.'
Best for: The default pick for coding agents, long-running agentic work and general knowledge work
2
Claude Fable 5.1 (Anthropic)
Released September 1, 2026, replacing Fable 5 in Anthropic's Mythos-class tier above Opus. Still $10 input / $50 output per million tokens, but cache reads drop to $0.25 per million, which Anthropic says cuts typical workload cost about 25% versus Fable 5. 1M-token context. #5 on LMArena at 1501. Scores 49.49% on Vals AI's Terminal-Bench 4.0, or 42.42% on the requested model alone (Anthropic reports 55.8%) and 65.0% on Humanity's Last Exam with tools.
Best for: The hardest reasoning and long-horizon agentic work, when Opus 5.5 still falls short on your evals
3
GPT-6 Astra (OpenAI)
OpenAI's new flagship, released September 3, 2026 at $10 input / $50 output per million tokens with a 1,050,000-token context window. Second on Vals AI's Terminal-Bench 4.0 at 57.07%, and first if fallback-served Claude tasks are counted as failures. OpenAI reports 99.9% on ARC-AGI-3 under its own harness, 97.6% on FrontierMath Tier 4 v2 and 72.6% on OSWorld 2.0 (offline set). It sits only #26 on LMArena at 1478; it is the one model placed here on Terminal-Bench 4.0 rather than LMArena order (see How We Ranked These).
Best for: Math, science and computer-use workloads, and the strongest non-Anthropic fallback
4
Muse Spark 1.3 (Meta)
Meta's current frontier model, which replaces open-weight Llama 4 as Meta's entry in this ranking. Shipped September 2, 2026 through the Meta Model API at $1.25 input / $4.25 output per million tokens, with a 1M-token context window. #9 on LMArena at 1494 (max reasoning). Meta reports 75.4% on DeepSWE v1.1 and 88.8% on Terminal-Bench 2.1 at max reasoning. Vals AI's Terminal-Bench 4.0 run scores the max variant 27.78%.
Best for: Near-top LMArena quality at roughly a tenth of Fable 5.1 or GPT-6 Astra output pricing
5
Gemini 3.8 Flash (Google)
Released September 2, 2026. Introductory pricing of $0.75 input / $3.75 output per million tokens runs through December 31, 2026, then rises to $1.50/$7.50. Input limit of 1,048,576 tokens. #10 on LMArena at 1492 (a score LMArena marks preliminary), the cheapest model in the top 10. Google reports 54.9% on HLE-Verified. Vals AI's Terminal-Bench 4.0 run scores it 13.13%.
Best for: High-volume coding and agent workloads where price matters more than the last few points of quality
6
Gemini 3.1 Pro (Google)
Still Google's newest 'Pro' tier, released February 19, 2026 and listed as a preview model, with a 1,048,576-token context window and 77.1% on ARC-AGI-2. #17 on LMArena at 1487, now behind Google's own Gemini 3.8 Flash. Google reports 80.6% on SWE-bench Verified, though independent runs land between 69.6% and 75.6%. Priced at $2/$12 under 200K tokens ($4/$18 above).
Best for: Long-document and multimodal work on Google Cloud at mid-tier pricing
7
Claude Sonnet 5 (Anthropic)
Released June 30, 2026 at $2 input / $10 output per million tokens, with a 1M-token context window. Scores 85.2% on SWE-bench Verified, 63.2% on SWE-bench Pro, and 80.4% on Terminal-Bench 2.1, a 13.4-point jump over the prior Sonnet 4.6, per Anthropic's system card. #54 on LMArena at 1462. Vals AI's Terminal-Bench 4.0 run scores it 8.08%, with seven provider refusals counted as failures.
Best for: Fast, everyday coding and knowledge work without Opus pricing
8
GPT-6 Sol (OpenAI)
Released September 22, 2026 alongside GPT-6 Luna, at $2 input / $10 output per million tokens (cached input $0.20) with a 1,050,000-token context window. That is half OpenAI's GPT-5.6 promotional pricing. OpenAI reports 33.2% on AutomationBench and 60.5% on OSWorld 2.0 at xhigh effort. #60 on LMArena at 1457. Terminal-Bench 4.0: not published.
Best for: Cost-sensitive business-workflow and computer-use automation inside the OpenAI stack
9
Grok 4.7 (xAI)
Released September 21, 2026 at Grok 4.6's unchanged $2 input / $6 output per million tokens (cached input $0.50; $4/$12 above 200K prompt tokens), with a 500,000-token context window. xAI reports 37.6% on Terminal-Bench 4.0 and 71.0% on DeepSWE v1.1 at high effort; Vals AI's independent Terminal-Bench 4.0 run scores it 28.28%. #92 on LMArena at 1439.
Best for: Low-cost xAI access at $2/$6, plus native real-time X data

Ranking order: LMArena text score (each model's best-scoring variant), with one exception: GPT-6 Astra is placed #3 on its Vals AI Terminal-Bench 4.0 result despite ranking #26 on LMArena. Sources for the ranking: LMArena ranks and scores from the LMArena text leaderboard (data dated September 25, 2026, read September 26). Independent Terminal-Bench 4.0 scores from Vals AI (updated September 22, 2026). Claude Opus 5.5 from Anthropic's Opus 5.5 announcement; Claude Fable 5.1 from Anthropic's Fable 5.1 page and MacRumors; Claude pricing and context from Anthropic's models overview. GPT-6 Astra pricing and context from OpenAI's model docs, release date from the GPT-6 Astra system card, OpenAI-reported benchmarks via DataCamp. GPT-6 Sol from OpenAI's model docs, TechCrunch and Vellum. Muse Spark 1.3 from Meta's developer page, release date from Meta's launch post. Gemini 3.8 Flash from Google's launch post and Gemini API docs. Grok 4.7 date, pricing and context from xAI's release notes, xAI-reported benchmarks from xAI's Grok 4.7 announcement.

Benchmark Scores Compared: Terminal-Bench 4.0

The September 2026 launches mostly stopped reporting SWE-bench Verified. Anthropic's Opus 5.5 announcement and Fable 5.1 page both report Terminal-Bench 4.0 and give no SWE-bench Verified figure, so this chart uses Terminal-Bench 4.0. It shows the only independent run we could find: Vals AI's Terminal-Bench 4.0 leaderboard, updated September 22, 2026. One caveat from Vals: 30 of Opus 5.5's 198 task attempts were served by Opus 5 or Opus 4.8 through provider-side fallback, and counting those as failures lowers it from 61.62% to 53.54%, behind GPT-6 Astra; Fable 5.1 scores 42.42% on the requested model alone. Vendor-run numbers are higher: Anthropic reports 66.4% for Opus 5.5, 55.8% for Fable 5.1 and 52.3% for Opus 5, alongside an OpenAI-reported 57.9% for GPT-6 Astra. Vals has not published a score for GPT-6 Sol, so it is left out of the chart rather than estimated.

On the older SWE-bench Verified benchmark, the highest published score among models in this post is still the superseded Claude Opus 5's 96.0%, followed by Claude Fable 5 at 95.0%, Claude Sonnet 5 at 85.2% (Anthropic's Claude Sonnet 5 system card), GPT-5.6 Sol at 82.2% and Gemini 3.1 Pro at a Google-reported 80.6%. None of the September 2026 releases in this ranking has a SWE-bench Verified figure in the sources we checked.

Which AI Model Is the Best Value in 2026: Price vs Performance

Price separates these models more than capability does. Output pricing runs from Gemini 3.8 Flash's introductory $3.75 per million tokens to $50 for Claude Fable 5.1 and GPT-6 Astra, a spread of more than 13x, while the LMArena gap between #1 and #10 is only 17 points. Here's the full side-by-side.

ModelInput / Output ($/1M tokens)LMArena text (rank)Terminal-Bench 4.0Context windowRank
Claude Opus 5.5$4 / $201509 (#1)61.6% Vals (53.5% excl. fallback); 66.4% Anthropic1M#1
Claude Fable 5.1$10 / $501501 (#5)49.5% Vals (42.4% excl. fallback); 55.8% Anthropic1M#2
GPT-6 Astra$10 / $50†1478 (#26)57.1% Vals1,050,000#3
Muse Spark 1.3$1.25 / $4.251494 (#9)27.8% Vals (max)1M#4
Gemini 3.8 Flash$0.75 / $3.75*1492 (#10, preliminary)13.1% Vals1,048,576#5
Gemini 3.1 Pro$2 / $12*1487 (#17)not published1,048,576#6
Claude Sonnet 5$2 / $101462 (#54)8.1% Vals1M#7
GPT-6 Sol$2 / $10†1457 (#60)not published1,050,000#8
Grok 4.7$2 / $6*1439 (#92)28.3% Vals; 37.6% xAI500K#9

*Gemini 3.8 Flash's $0.75/$3.75 is introductory through December 31, 2026, then $1.50/$7.50 (Gemini API pricing). Gemini 3.1 Pro rises to $4/$18 above 200K tokens, and Grok 4.7 bills $4/$12 above 200K prompt tokens. †OpenAI prices GPT-6 prompts above 272K input tokens at 2x input and 1.5x output for the full request. LMArena scores were read on September 26, 2026 from board data dated September 25; Vals AI Terminal-Bench 4.0 scores are as of September 22, 2026. Pricing and benchmarks change frequently, so verify current figures before committing production spend.

Superseded Since the August 31 Edition

Four models from the previous edition have newer versions from the same vendor. Claude Fable 5 (formerly #1) is replaced by Fable 5.1 but remains available as a legacy model and still sits #3 on LMArena at 1504. Claude Opus 5 (formerly #2) is replaced by Opus 5.5, which Anthropic prices 20% lower. GPT-5.6 Sol (formerly #3, #19 on LMArena at 1483) is replaced by GPT-6 Sol at half its promotional price, with GPT-6 Astra above it. Grok 4.6 (formerly #6) is replaced by Grok 4.7 at the same price. Meta's open-weight Llama 4 drops out of the ranking: Muse Spark 1.3 is now Meta's frontier model, though Llama 4 remains the option for teams that need to self-host weights with zero per-token licensing cost.

How We Ranked These

The order starts from the LMArena text leaderboard as read on September 26, 2026 (board data dated September 25, 2026), counting each model's best-scoring variant. We made one deliberate exception: GPT-6 Astra sits #26 on LMArena, but ranks third here because it is second on Vals AI's independent Terminal-Bench 4.0 run (first if fallback-served Claude tasks are counted as failures) and is OpenAI's flagship. Every other model follows its LMArena order. This is an editorial blend, not a formula, so the inputs are published above for you to re-weight. Vendor-reported benchmarks are labelled as such, and we used independent runs from Vals AI wherever one existed, because self-reported and independently run numbers diverge. Google's own 80.6% SWE-bench Verified claim for Gemini 3.1 Pro comes down to 69.6-75.6% in third-party runs, and Anthropic's 66.4% Terminal-Bench 4.0 figure for Opus 5.5 comes down to 61.6% in Vals AI's run. Where no source published a number, the table says "not published" instead of estimating one. Pricing was checked against each vendor's own API documentation in September 2026. Sponsors and affiliate partners never influence rank or inclusion — see our editorial standards.

Why the AI Model Rankings Keep Shifting in 2026

Three years ago, a new frontier model release moved the leaderboard for months. In September 2026 alone, Anthropic shipped Claude Fable 5.1 (September 1) and Claude Opus 5.5 (September 22). OpenAI shipped GPT-6 Astra (September 3) and then GPT-6 Sol and Luna (September 22), less than three months after the GPT-5.6 family. Google shipped Gemini 3.8 Flash (September 2), Meta shipped Muse Spark 1.3 the same day, and xAI shipped Grok 4.7 on September 21. Google still has no Pro model newer than Gemini 3.1 Pro, which is why a Flash model now outranks it on LMArena. That pace matters more than any single benchmark number: picking a model on brand reputation alone, rather than task-specific performance and current pricing, now leaves real accuracy and cost on the table.

It also means the "best model" question has split into separate questions: best for coding, best for autonomous agents, best for cost and best overall reasoning. Claude Opus 5.5 currently leads both LMArena and the headline independent Terminal-Bench 4.0 score, which is unusual, though its Terminal-Bench lead depends on how fallback-served tasks are counted. GPT-6 Astra leads the math and ARC-AGI-3 results OpenAI published, and the cheapest options sit near the top of LMArena but score far lower on independent Terminal-Bench 4.0: 27.8% for Muse Spark 1.3 and 13.1% for Gemini 3.8 Flash. Treat each benchmark as measuring a genuinely different capability, not a proxy for overall quality.

What the headline misses

A benchmark table implies these models are equally available to build on, and they aren't always. Claude Fable 5, the model that topped this ranking until this update, was unreachable worldwide for about 19 days in June 2026 after a US export-control directive suspended it, with no public warning before the suspension or guaranteed timeline for the redeployment that followed. GPT-6 Astra also rolled out to a limited set of organizations first. A team that had shipped a production dependency on Fable 5 in that window would have had to fail over to a different vendor with zero notice. This likely means the practical "best model" decision for anything customer-facing should weight platform stability and a credible fallback path at least as heavily as the top-line benchmark score, something none of the leaderboards cited above actually measure.

How to Choose the Best AI Model for Your Use Case

Start with the task, not the vendor. For coding agents and long-running agentic work, Claude Opus 5.5 is the strongest pick: it has the top independent Terminal-Bench 4.0 score at $4/$20 per million tokens. Step up to Claude Fable 5.1 only if your own evals show Opus 5.5 falling short. For math, science or computer-use workloads, or if you need a non-Anthropic primary, GPT-6 Astra is the strongest alternative at the same $10/$50 as Fable 5.1. For high-volume work where frontier reasoning is overkill, Gemini 3.8 Flash at an introductory $0.75/$3.75 and Muse Spark 1.3 at $1.25/$4.25 both sit in LMArena's top 10 at a fraction of flagship prices. Budget for Gemini 3.8 Flash's price doubling on January 1, 2027.

For portfolio companies still deciding between a single-vendor commitment and a routing layer, I recommend routing by default once monthly API spend crosses roughly $5,000/month. Below that threshold, engineering time spent building a router usually costs more than the token savings justify. Above it, a lightweight router that sends coding and agent tasks to Opus 5.5 and bulk tasks to Gemini 3.8 Flash or GPT-6 Luna, with GPT-6 Astra as a documented fallback, cuts blended cost versus single-vendor lock-in. The fallback also protects against a repeat of the Fable 5 export-control gap.

Track how these model economics flow into startup valuations on the AI Valuations Dashboard and how the hyperscalers funding this buildout are reporting results on Big Tech Earnings at Value Add VC.

What AI Model Benchmarks Actually Measure

Terminal-Bench 4.0, the coding benchmark behind the chart above, measures whether a model can operate autonomously inside a command-line environment across many sequential steps without a human checking each one. That means installing dependencies, debugging a failing build and recovering from its own mistakes. Scores are far lower than on older benchmarks because the tasks are harder: even the leader, Claude Opus 5.5, completes about 62% in Vals AI's run.

SWE-bench Verified, which drove this ranking's earlier editions, tests models against real GitHub issues pulled from popular open-source repositories. The model has to read the codebase, understand the bug report and produce a patch that passes the project's actual test suite. With the best published scores at 95-96%, it no longer separates frontier models, which is one reason the September 2026 launches reported Terminal-Bench 4.0 instead.

LMArena's score works differently again. It is a crowdsourced head-to-head preference ranking, not a task-completion score, generated from blind pairwise comparisons submitted by real users. It captures something closer to "which answer do people prefer reading," which is why GPT-6 Astra can be second on Terminal-Bench 4.0 while sitting #26 on LMArena. Use LMArena as a general-quality signal and the task-specific benchmarks as the actual purchasing decision.

The Bottom Line

There is no single best AI model for every job in 2026, but as of September 26 one model leads on both crowd preference and the headline independent agentic-coding score.

Claude Opus 5.5 wins the overall crown, Claude Fable 5.1 and GPT-6 Astra are the premium alternatives, Gemini 3.8 Flash wins on price — pick by workload, not by brand loyalty.

Track AI company valuations and the hyperscaler capex funding this model race on the AI Valuations Dashboard at Value Add VC. Originally published in the Trace Cohen newsletter.

Get VC data most people never see

— free to subscribe

Trace's notes on venture, AI and startups, a few times a week. Join 5,000+ subscribers. No spam.

Frequently Asked Questions

What is the best AI model in September 2026?

Claude Opus 5.5 from Anthropic, released September 22, 2026. It sits #1 on LMArena's text leaderboard at 1509 (board data dated September 25, read September 26) and leads Vals AI's independent Terminal-Bench 4.0 run at 61.6%, ahead of GPT-6 Astra at 57.1% and Claude Fable 5.1 at 49.5%, though Vals notes that 30 of Opus 5.5's 198 task attempts were served by older Claude models; counting those as failures drops it to 53.5%, behind Astra. Anthropic says it performs at the level of Fable 5.1 on most work while costing $4/$20 per million tokens instead of $10/$50.

Which AI model is cheapest at the frontier in 2026?

Among models in LMArena's top 10, Google's Gemini 3.8 Flash is the cheapest at an introductory $0.75 input / $3.75 output per million tokens through December 31, 2026, rising to $1.50/$7.50 on January 1, 2027. Meta's Muse Spark 1.3 is next at $1.25/$4.25. Further down the board, OpenAI's GPT-6 Luna lists at $0.10/$0.50 and GPT-6 Sol at $2/$10, while Grok 4.7 keeps Grok 4.6's $2/$6 pricing.

Is GPT-6 Astra or Claude Opus 5.5 better for coding?

It is close as of September 26, 2026. Vals AI's independent Terminal-Bench 4.0 run has Opus 5.5 at 61.6% versus 57.1% for GPT-6 Astra, but Vals notes 30 of Opus 5.5's 198 task attempts were served by older Claude models; counted as failures, Opus 5.5 falls to 53.5%, behind Astra. Anthropic's table shows 66.4% for Opus 5.5 against an OpenAI-reported 57.9% for Astra. Opus 5.5 costs less, at $4/$20 per million tokens versus Astra's $10/$50. OpenAI also reports 99.9% for Astra on ARC-AGI-3 under its own harness and 97.6% on FrontierMath Tier 4 v2.

What happened to Claude Fable 5's availability in June 2026?

Anthropic released Claude Fable 5 on June 9, 2026 as its first public 'Mythos-class' model, a tier positioned above Opus. Three days later, on June 12, 2026, Anthropic received a US government export-control directive requiring it to suspend access to both Fable 5 and the unreleased Mythos 5 model worldwide. The directive was lifted on July 1, 2026, and Anthropic began redeploying Fable 5 globally that day. Fable 5 has since been superseded by Claude Fable 5.1, released September 1, 2026.

Which AI model has the largest context window in 2026?

Most of the frontier now sits at about 1 million tokens. OpenAI lists a 1,050,000-token context window for GPT-6 Astra, Sol and Luna; Anthropic lists 1M tokens for Claude Opus 5.5, Fable 5.1 and Sonnet 5; Google lists a 1,048,576-token input limit for Gemini 3.8 Flash; and Meta lists 1M for Muse Spark 1.3. Grok 4.7 is the outlier at 500,000 tokens. Meta's older open-weight Llama 4 Scout still advertises 10 million tokens.

Should a startup use one AI model or multiple in 2026?

Most production teams getting the best results in 2026 run a routing layer rather than committing to a single vendor. A typical September 2026 split: Claude Opus 5.5 for coding and long-running agents, Gemini 3.8 Flash or GPT-6 Luna for high-volume work where frontier reasoning is overkill, and a second vendor such as GPT-6 Astra as a documented fallback. With Anthropic, OpenAI and Google each shipping at least one new model in September alone, single-vendor lock-in leaves both cost and accuracy on the table.

Explore 45+ free VC tools, dashboards, and recommended startup software.

Get VC data most people never see

Venture, AI & startup notes a few times a week. Join 5,000+ subscribers.