Claude Opus 5.5 is the best AI model as of September 26, 2026: #1 on LMArena at 1509 and #1 on an independent Terminal-Bench 4.0 run at 61.6%, at $4/$20 per million tokens. That is less than half the list price of Claude Fable 5.1 or GPT-6 Astra, the two models just behind it. Nine models make this list, and picking the wrong one for the job, not the wrong vendor, is where most teams now lose money.
I evaluate AI tools across 65+ portfolio companies, and September 2026 was the busiest month for frontier releases this year. In four weeks Anthropic shipped Claude Fable 5.1 and Claude Opus 5.5, OpenAI shipped GPT-6 Astra followed by GPT-6 Sol and Luna, Google shipped Gemini 3.8 Flash, Meta shipped Muse Spark 1.3, and xAI shipped Grok 4.7. Four of the seven models in this ranking's August 31 edition have been superseded.

What Are the Best AI Models in September 2026, Ranked?
Claude Opus 5.5 from Anthropic is the top-ranked AI model as of September 26, 2026. It leads the LMArena text leaderboard at 1509 (board data dated September 25, 2026, read September 26) and the Vals AI Terminal-Bench 4.0 leaderboard at 61.62% (updated September 22, 2026). Claude Fable 5.1 ranks second and GPT-6 Astra third. Gemini 3.8 Flash is the cheapest model in LMArena's top 10.
Meta's Muse Spark 1.3, Google's Gemini 3.1 Pro, Claude Sonnet 5, OpenAI's GPT-6 Sol and xAI's Grok 4.7 round out the nine. The top of LMArena is tight: Opus 5.5's 1509 carries a ±12 confidence interval, and the superseded Claude Fable 5 still sits #3 at 1504. Treat a gap of a few points as a tie. Track the category leader for your specific task, not the vendor brand, when picking a model for a new build.
The 9 Best AI Models in 2026, Ranked (as of September 26)
Ranking order: LMArena text score (each model's best-scoring variant), with one exception: GPT-6 Astra is placed #3 on its Vals AI Terminal-Bench 4.0 result despite ranking #26 on LMArena. Sources for the ranking: LMArena ranks and scores from the LMArena text leaderboard (data dated September 25, 2026, read September 26). Independent Terminal-Bench 4.0 scores from Vals AI (updated September 22, 2026). Claude Opus 5.5 from Anthropic's Opus 5.5 announcement; Claude Fable 5.1 from Anthropic's Fable 5.1 page and MacRumors; Claude pricing and context from Anthropic's models overview. GPT-6 Astra pricing and context from OpenAI's model docs, release date from the GPT-6 Astra system card, OpenAI-reported benchmarks via DataCamp. GPT-6 Sol from OpenAI's model docs, TechCrunch and Vellum. Muse Spark 1.3 from Meta's developer page, release date from Meta's launch post. Gemini 3.8 Flash from Google's launch post and Gemini API docs. Grok 4.7 date, pricing and context from xAI's release notes, xAI-reported benchmarks from xAI's Grok 4.7 announcement.
Benchmark Scores Compared: Terminal-Bench 4.0
The September 2026 launches mostly stopped reporting SWE-bench Verified. Anthropic's Opus 5.5 announcement and Fable 5.1 page both report Terminal-Bench 4.0 and give no SWE-bench Verified figure, so this chart uses Terminal-Bench 4.0. It shows the only independent run we could find: Vals AI's Terminal-Bench 4.0 leaderboard, updated September 22, 2026. One caveat from Vals: 30 of Opus 5.5's 198 task attempts were served by Opus 5 or Opus 4.8 through provider-side fallback, and counting those as failures lowers it from 61.62% to 53.54%, behind GPT-6 Astra; Fable 5.1 scores 42.42% on the requested model alone. Vendor-run numbers are higher: Anthropic reports 66.4% for Opus 5.5, 55.8% for Fable 5.1 and 52.3% for Opus 5, alongside an OpenAI-reported 57.9% for GPT-6 Astra. Vals has not published a score for GPT-6 Sol, so it is left out of the chart rather than estimated.
On the older SWE-bench Verified benchmark, the highest published score among models in this post is still the superseded Claude Opus 5's 96.0%, followed by Claude Fable 5 at 95.0%, Claude Sonnet 5 at 85.2% (Anthropic's Claude Sonnet 5 system card), GPT-5.6 Sol at 82.2% and Gemini 3.1 Pro at a Google-reported 80.6%. None of the September 2026 releases in this ranking has a SWE-bench Verified figure in the sources we checked.
Which AI Model Is the Best Value in 2026: Price vs Performance
Price separates these models more than capability does. Output pricing runs from Gemini 3.8 Flash's introductory $3.75 per million tokens to $50 for Claude Fable 5.1 and GPT-6 Astra, a spread of more than 13x, while the LMArena gap between #1 and #10 is only 17 points. Here's the full side-by-side.
| Model | Input / Output ($/1M tokens) | LMArena text (rank) | Terminal-Bench 4.0 | Context window | Rank |
|---|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | 1509 (#1) | 61.6% Vals (53.5% excl. fallback); 66.4% Anthropic | 1M | #1 |
| Claude Fable 5.1 | $10 / $50 | 1501 (#5) | 49.5% Vals (42.4% excl. fallback); 55.8% Anthropic | 1M | #2 |
| GPT-6 Astra | $10 / $50† | 1478 (#26) | 57.1% Vals | 1,050,000 | #3 |
| Muse Spark 1.3 | $1.25 / $4.25 | 1494 (#9) | 27.8% Vals (max) | 1M | #4 |
| Gemini 3.8 Flash | $0.75 / $3.75* | 1492 (#10, preliminary) | 13.1% Vals | 1,048,576 | #5 |
| Gemini 3.1 Pro | $2 / $12* | 1487 (#17) | not published | 1,048,576 | #6 |
| Claude Sonnet 5 | $2 / $10 | 1462 (#54) | 8.1% Vals | 1M | #7 |
| GPT-6 Sol | $2 / $10† | 1457 (#60) | not published | 1,050,000 | #8 |
| Grok 4.7 | $2 / $6* | 1439 (#92) | 28.3% Vals; 37.6% xAI | 500K | #9 |
*Gemini 3.8 Flash's $0.75/$3.75 is introductory through December 31, 2026, then $1.50/$7.50 (Gemini API pricing). Gemini 3.1 Pro rises to $4/$18 above 200K tokens, and Grok 4.7 bills $4/$12 above 200K prompt tokens. †OpenAI prices GPT-6 prompts above 272K input tokens at 2x input and 1.5x output for the full request. LMArena scores were read on September 26, 2026 from board data dated September 25; Vals AI Terminal-Bench 4.0 scores are as of September 22, 2026. Pricing and benchmarks change frequently, so verify current figures before committing production spend.
Superseded Since the August 31 Edition
Four models from the previous edition have newer versions from the same vendor. Claude Fable 5 (formerly #1) is replaced by Fable 5.1 but remains available as a legacy model and still sits #3 on LMArena at 1504. Claude Opus 5 (formerly #2) is replaced by Opus 5.5, which Anthropic prices 20% lower. GPT-5.6 Sol (formerly #3, #19 on LMArena at 1483) is replaced by GPT-6 Sol at half its promotional price, with GPT-6 Astra above it. Grok 4.6 (formerly #6) is replaced by Grok 4.7 at the same price. Meta's open-weight Llama 4 drops out of the ranking: Muse Spark 1.3 is now Meta's frontier model, though Llama 4 remains the option for teams that need to self-host weights with zero per-token licensing cost.
How We Ranked These
The order starts from the LMArena text leaderboard as read on September 26, 2026 (board data dated September 25, 2026), counting each model's best-scoring variant. We made one deliberate exception: GPT-6 Astra sits #26 on LMArena, but ranks third here because it is second on Vals AI's independent Terminal-Bench 4.0 run (first if fallback-served Claude tasks are counted as failures) and is OpenAI's flagship. Every other model follows its LMArena order. This is an editorial blend, not a formula, so the inputs are published above for you to re-weight. Vendor-reported benchmarks are labelled as such, and we used independent runs from Vals AI wherever one existed, because self-reported and independently run numbers diverge. Google's own 80.6% SWE-bench Verified claim for Gemini 3.1 Pro comes down to 69.6-75.6% in third-party runs, and Anthropic's 66.4% Terminal-Bench 4.0 figure for Opus 5.5 comes down to 61.6% in Vals AI's run. Where no source published a number, the table says "not published" instead of estimating one. Pricing was checked against each vendor's own API documentation in September 2026. Sponsors and affiliate partners never influence rank or inclusion — see our editorial standards.
Why the AI Model Rankings Keep Shifting in 2026
Three years ago, a new frontier model release moved the leaderboard for months. In September 2026 alone, Anthropic shipped Claude Fable 5.1 (September 1) and Claude Opus 5.5 (September 22). OpenAI shipped GPT-6 Astra (September 3) and then GPT-6 Sol and Luna (September 22), less than three months after the GPT-5.6 family. Google shipped Gemini 3.8 Flash (September 2), Meta shipped Muse Spark 1.3 the same day, and xAI shipped Grok 4.7 on September 21. Google still has no Pro model newer than Gemini 3.1 Pro, which is why a Flash model now outranks it on LMArena. That pace matters more than any single benchmark number: picking a model on brand reputation alone, rather than task-specific performance and current pricing, now leaves real accuracy and cost on the table.
It also means the "best model" question has split into separate questions: best for coding, best for autonomous agents, best for cost and best overall reasoning. Claude Opus 5.5 currently leads both LMArena and the headline independent Terminal-Bench 4.0 score, which is unusual, though its Terminal-Bench lead depends on how fallback-served tasks are counted. GPT-6 Astra leads the math and ARC-AGI-3 results OpenAI published, and the cheapest options sit near the top of LMArena but score far lower on independent Terminal-Bench 4.0: 27.8% for Muse Spark 1.3 and 13.1% for Gemini 3.8 Flash. Treat each benchmark as measuring a genuinely different capability, not a proxy for overall quality.
What the headline misses
A benchmark table implies these models are equally available to build on, and they aren't always. Claude Fable 5, the model that topped this ranking until this update, was unreachable worldwide for about 19 days in June 2026 after a US export-control directive suspended it, with no public warning before the suspension or guaranteed timeline for the redeployment that followed. GPT-6 Astra also rolled out to a limited set of organizations first. A team that had shipped a production dependency on Fable 5 in that window would have had to fail over to a different vendor with zero notice. This likely means the practical "best model" decision for anything customer-facing should weight platform stability and a credible fallback path at least as heavily as the top-line benchmark score, something none of the leaderboards cited above actually measure.
How to Choose the Best AI Model for Your Use Case
Start with the task, not the vendor. For coding agents and long-running agentic work, Claude Opus 5.5 is the strongest pick: it has the top independent Terminal-Bench 4.0 score at $4/$20 per million tokens. Step up to Claude Fable 5.1 only if your own evals show Opus 5.5 falling short. For math, science or computer-use workloads, or if you need a non-Anthropic primary, GPT-6 Astra is the strongest alternative at the same $10/$50 as Fable 5.1. For high-volume work where frontier reasoning is overkill, Gemini 3.8 Flash at an introductory $0.75/$3.75 and Muse Spark 1.3 at $1.25/$4.25 both sit in LMArena's top 10 at a fraction of flagship prices. Budget for Gemini 3.8 Flash's price doubling on January 1, 2027.
For portfolio companies still deciding between a single-vendor commitment and a routing layer, I recommend routing by default once monthly API spend crosses roughly $5,000/month. Below that threshold, engineering time spent building a router usually costs more than the token savings justify. Above it, a lightweight router that sends coding and agent tasks to Opus 5.5 and bulk tasks to Gemini 3.8 Flash or GPT-6 Luna, with GPT-6 Astra as a documented fallback, cuts blended cost versus single-vendor lock-in. The fallback also protects against a repeat of the Fable 5 export-control gap.
Track how these model economics flow into startup valuations on the AI Valuations Dashboard and how the hyperscalers funding this buildout are reporting results on Big Tech Earnings at Value Add VC.
What AI Model Benchmarks Actually Measure
Terminal-Bench 4.0, the coding benchmark behind the chart above, measures whether a model can operate autonomously inside a command-line environment across many sequential steps without a human checking each one. That means installing dependencies, debugging a failing build and recovering from its own mistakes. Scores are far lower than on older benchmarks because the tasks are harder: even the leader, Claude Opus 5.5, completes about 62% in Vals AI's run.
SWE-bench Verified, which drove this ranking's earlier editions, tests models against real GitHub issues pulled from popular open-source repositories. The model has to read the codebase, understand the bug report and produce a patch that passes the project's actual test suite. With the best published scores at 95-96%, it no longer separates frontier models, which is one reason the September 2026 launches reported Terminal-Bench 4.0 instead.
LMArena's score works differently again. It is a crowdsourced head-to-head preference ranking, not a task-completion score, generated from blind pairwise comparisons submitted by real users. It captures something closer to "which answer do people prefer reading," which is why GPT-6 Astra can be second on Terminal-Bench 4.0 while sitting #26 on LMArena. Use LMArena as a general-quality signal and the task-specific benchmarks as the actual purchasing decision.
The Bottom Line
There is no single best AI model for every job in 2026, but as of September 26 one model leads on both crowd preference and the headline independent agentic-coding score.
Claude Opus 5.5 wins the overall crown, Claude Fable 5.1 and GPT-6 Astra are the premium alternatives, Gemini 3.8 Flash wins on price — pick by workload, not by brand loyalty.
Track AI company valuations and the hyperscaler capex funding this model race on the AI Valuations Dashboard at Value Add VC. Originally published in the Trace Cohen newsletter.
Latest from the Pulse
Get VC data most people never see
— free to subscribe
Trace's notes on venture, AI and startups, a few times a week. Join 5,000+ subscribers. No spam.