Which AI model is best at which PM task?

The only AI benchmark that matters to product execs. Frontier LLMs tested to the standards of world-class product and growth leaders.

Very cool! Love this.
Lenny Rachitsky

8 models23 task types368 outputs graded

GPT-6 Astra leads by 2 points.

Combined score by model and task type. Choose a column heading to sort by that task.
ModelProduct CraftGrowthMetrics & ExperimentationStrategyLeadership
GPT-6 AstraOpenAI90#196Best8184Best97708991Best84100Best98Best7998Best9199Best87Best88929396Best8694Best8691Best
GPT-6.1 SolOpenAI88#295857198Best899189878497758396997791899380818896Best89
Sonnet 5.5Anthropic87#38086Best779493Best94Best9090Best848895Best9296Best92787881838584909281
Haiku 5.5Anthropic84#47777678692928270829689788696809293Best808176897991
Opus 5.5Anthropic82#5728265947791827291928184868271628494Best8289Best818387
GPT-6 LunaOpenAI81#687776587857786669098767592827493Best84867968718282
Gemini 3.8 FlashGoogle61#76658396876555661777160557161615043606351586982
Gemini 3.5 Flash-LiteGoogle49#85641394449524435687429534451495040415642606061
Combined scoreEach cell is the model’s best setup on that task type.

Tested with criteria sourced from the very best.

We carefully research and design the rubric for each task around the advice of guests on Lenny’s Podcast.

30 product leaders behind the checks

We credit their ideas. None of them has reviewed or endorsed the benchmark.

Inside the benchmark

The checks AI fails most

  1. Names an advantage that's hard to copyDevelop product strategy11%
  2. Marks what to cut if the call runs overCustomer research call guide28%
  3. Doesn't trust the default reasonCustomer research call guide41%
  4. Says what it tests and what's fakedOne-shot prototype44%
What AI gets right and wrong, task by task

What you get for the cost

Show

GPT-6.1 Sol · API scores within 2 points of GPT-6 Astra · ChatGPT at a fifth of the cost.

Task score against relative cost for each setup. The table below has the same figures.4050607080901001×3×10×30×Cost per output, relative to the cheapest setup (log scale) →↑ Task scoreGPT-6 Astra · ChatGPT: task score 90, 68× the cheapest (10 of 46 outputs estimated)GPT-6.1 Sol · API: task score 88, 14× the cheapestSonnet 5.5 · API: task score 87, 44× the cheapestHaiku 5.5 · API: task score 84, 2.4× the cheapestOpus 5.5 · Claude: task score 82, 65× the cheapest (10 of 46 outputs estimated), 1 critical failureGPT-6 Luna · API: task score 81, the cheapestGemini 3.8 Flash · API: task score 61, 9.9× the cheapest, 2 critical failuresGemini 3.5 Flash-Lite · Gemini: task score 49, 2.2× the cheapest (11 of 46 outputs estimated), 5 critical failuresGPT-6 Astra · ChatGPTGPT-6.1 Sol · APISonnet 5.5 · APIHaiku 5.5 · APIOpus 5.5 · ClaudeGPT-6 Luna · APIGemini 3.8 Flash · APIGemini 3.5 Flash-Lite · Gemini
Best valueNobody beats these on both score and costCost partly estimated (app runs with no usage meter)At least one critical failure
Task score and relative cost by setup
SetupTask scoreRelative cost (cheapest = 1)
GPT-6 Luna · API811.0×
Gemini 3.5 Flash-Lite · Gemini492.2× (partly estimated)
Haiku 5.5 · API842.4×
Gemini 3.8 Flash · API619.9×
GPT-6.1 Sol · API8814×
Sonnet 5.5 · API8744×
Opus 5.5 · Claude8265× (partly estimated)
GPT-6 Astra · ChatGPT9068× (partly estimated)

Score is the Combined score across every task type; cost is each setup’s typical output, relative to the cheapest.How we compare cost

Recent drops

All drops