VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog๐ŸคPartner
Illustration for: OpenAI Releases LifeSciBench -- Frontier Models Pass Just 1 in 3 Real Science Tasks
Value Add VC/Pulse/AI

OpenAI Releases LifeSciBench -- Frontier Models Pass Just 1 in 3 Real Science Tasks

OpenAI introduced LifeSciBench, an expert-authored benchmark of 750 tasks across seven biological domains, built with 173 scientists. Even the strongest models pass only about one task in three -- a sober counterweight to claims that AI is close to autonomous scientific research.

By the Numbers

750
Tasks
173
Contributing Scientists
~33%
Top Model Pass Rate
7
Domains
TC
Trace Cohen
Early-stage VC & angel ยท Founder, New York Venture Partners
June 17, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

A credible ceiling on AI-for-science hype: today's frontier still fails two of three real research tasks

2

Expert-written rubrics from 173 scientists make this a hard benchmark to game -- the gap is real, not artifactual

3

For founders selling 'AI scientist' products, this is the honest baseline buyers will now measure against

TC

The VC Read ยท Trace's Take

Trace Cohen

OpenAI publishing a benchmark its own models fail two-thirds of is the most useful thing it shipped this week. The 'AI scientist' pitch has gotten ahead of reality, and now there's a hard, expert-graded number to anchor the conversation. For founders selling autonomous research, this is the baseline your buyers will quote back at you -- so position as augmentation, not replacement, or get caught overclaiming. The honest framing is also the more defensible business.

๐Ÿค– AI Landscape โ†’๐Ÿ“Š Benchmarking โ†’

Analysis

OpenAI released LifeSciBench, a benchmark of 750 expert-authored tasks spanning seven biological domains and seven research workflows, developed with input from 173 scientists. The headline result is humbling: even the most capable frontier models pass only roughly one in three tasks, with detailed rubrics grading the quality of scientific reasoning rather than just final answers.

The benchmark matters because it pushes back against the loudest version of the AI-for-science narrative. While models can accelerate literature review, hypothesis generation, and some analysis, LifeSciBench shows they remain far from autonomously executing real research workflows. The expert-written rubrics make the result hard to dismiss as a measurement artifact.

โ€œThe benchmark matters because it pushes back against the loudest version of the AI-for-science narrative.โ€

For builders, the takeaway is to calibrate. AI is a powerful copilot for scientists, not a replacement, and the honest framing -- augment the researcher, don't replace them -- is both more defensible commercially and more credible with the domain buyers who will increasingly benchmark these claims against tools like LifeSciBench.

ShareXLinkedInEmail

More on

OpenAI โ†’

Analysis and editorial commentary by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AIยท Aug 3, 2026

DeepMind Exec: AI Capex Is Really a Bet on Self-Improvement

Illustration for: DeepMind Exec: AI Capex Is Really a Bet on Self-Improvement
AI

DeepMind Exec: AI Capex Is Really a Bet on Self-Improvement

A Google DeepMind executive says unprecedented data-center spending only makes sense as a bet AI systems will soon meaningfully improve themselves -- a far more aggressive rationale than usual capex talk.

AIยท Aug 3, 2026

Alibaba Undercuts Kimi K3 With a Cheaper Flagship Model

Illustration for: Alibaba Undercuts Kimi K3 With a Cheaper Flagship Model
AI

Alibaba Undercuts Kimi K3 With a Cheaper Flagship Model

Alibaba released a new flagship model priced below Kimi K3, its latest open-weight swipe at US frontier labs, extending a pattern of Chinese labs competing primarily on price rather than raw benchmark scores.

AIยท Aug 2, 2026

DeepSeek's Cheap New Model Is Making Waves Again

Illustration for: DeepSeek's Cheap New Model Is Making Waves Again
AI

DeepSeek's Cheap New Model Is Making Waves Again

DeepSeek's new V4-Flash model is drawing outsized attention for its price-to-performance ratio, reigniting the cost-disruption dynamic that made DeepSeek a household name earlier this year.

@Trace_Cohenยทt@nyvp.com