Which AI model is best at which PM task?
The only AI benchmark that matters to product execs. Frontier LLMs tested to the standards of world-class product and growth leaders.
Very cool! Love this.
8 models23 task types368 outputs graded
GPT-6 Astra leads by 2 points.
| Model | Product Craft | Growth | Metrics & Experimentation | Strategy | Leadership | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 90#1 | 96Best | 81 | 84Best | 97 | 70 | 89 | 91Best | 84 | 100Best | 98Best | 79 | 98Best | 91 | 99Best | 87Best | 88 | 92 | 93 | 96Best | 86 | 94Best | 86 | 91Best |
| GPT-6.1 SolOpenAI | 88#2 | 95 | 85 | 71 | 98Best | 89 | 91 | 89 | 87 | 84 | 97 | 75 | 83 | 96 | 99 | 77 | 91 | 89 | 93 | 80 | 81 | 88 | 96Best | 89 |
| Sonnet 5.5Anthropic | 87#3 | 80 | 86Best | 77 | 94 | 93Best | 94Best | 90 | 90Best | 84 | 88 | 95Best | 92 | 96Best | 92 | 78 | 78 | 81 | 83 | 85 | 84 | 90 | 92 | 81 |
| Haiku 5.5Anthropic | 84#4 | 77 | 77 | 67 | 86 | 92 | 92 | 82 | 70 | 82 | 96 | 89 | 78 | 86 | 96 | 80 | 92 | 93Best | 80 | 81 | 76 | 89 | 79 | 91 |
| Opus 5.5Anthropic | 82#5 | 72 | 82 | 65 | 94 | 77 | 91 | 82 | 72 | 91 | 92 | 81 | 84 | 86 | 82 | 71 | 62 | 84 | 94Best | 82 | 89Best | 81 | 83 | 87 |
| GPT-6 LunaOpenAI | 81#6 | 87 | 77 | 65 | 87 | 85 | 77 | 86 | 66 | 90 | 98 | 76 | 75 | 92 | 82 | 74 | 93Best | 84 | 86 | 79 | 68 | 71 | 82 | 82 |
| Gemini 3.8 FlashGoogle | 61#7 | 66 | 58 | 39 | 68 | 76 | 55 | 56 | 61 | 77 | 71 | 60 | 55 | 71 | 61 | 61 | 50 | 43 | 60 | 63 | 51 | 58 | 69 | 82 |
| Gemini 3.5 Flash-LiteGoogle | 49#8 | 56 | 41 | 39 | 44 | 49 | 52 | 44 | 35 | 68 | 74 | 29 | 53 | 44 | 51 | 49 | 50 | 40 | 41 | 56 | 42 | 60 | 60 | 61 |
Tested with criteria sourced from the very best.
We carefully research and design the rubric for each task around the advice of guests on Lenny’s Podcast.
30 product leaders behind the checks
- Lauryn IsfordActivation +2
- Bangaly KabaActivation +2
- Hamilton HelmerStrategy +2
- Matt LeMayPRD +1
- April DunfordGTM plan +1
- Christian IdiodiRoadmap +1
- Judd AntinDiscovery +1
- Melissa PerriRoadmap +1
- Ami VoraPRD
- Christina WodtkeOKRs
- Emily KramerGTM plan
- Hila QuActivation
- Jiaona ZhangRoadmap
- Naomi GleitRetention
- Rahul VohraPMF call
- Ronny KohaviTest design +1
- Gokul RajaramRoadmap +2
- Karri SaarinenPRD +2
- Todd JacksonLaunch call +1
- Bob MoestaDiscovery +1
- Itamar GiladNorth Star
- Madhavan RamanujamPricing
- Sean EllisNorth Star +1
- Arielle JacksonGTM plan
- Claire Hughes JohnsonOKRs
- Geoff CharlesExec update
- Jeanne DeWitt GrosserGTM plan
- Kim ScottReviews
- Petra WilleHiring
- Teresa TorresDiscovery
We credit their ideas. None of them has reviewed or endorsed the benchmark.
Inside the benchmark
The mistakes PMs fix most
Every mistake, with real examplesWhat you get for the cost
GPT-6.1 Sol · API scores within 2 points of GPT-6 Astra · ChatGPT at a fifth of the cost.
| Setup | Task score | Relative cost (cheapest = 1) |
|---|---|---|
| GPT-6 Luna · API | 81 | 1.0× |
| Gemini 3.5 Flash-Lite · Gemini | 49 | 2.2× (partly estimated) |
| Haiku 5.5 · API | 84 | 2.4× |
| Gemini 3.8 Flash · API | 61 | 9.9× |
| GPT-6.1 Sol · API | 88 | 14× |
| Sonnet 5.5 · API | 87 | 44× |
| Opus 5.5 · Claude | 82 | 65× (partly estimated) |
| GPT-6 Astra · ChatGPT | 90 | 68× (partly estimated) |
Score is the Combined score across every task type; cost is each setup’s typical output, relative to the cheapest.How we compare cost
Recent drops
Claude Haiku 5.5
Nearly Sonnet 5.5's judgement. Watch it with numbers & calculations.
What you can hand it, and what you’ll still have to catchGemini 3.8 Flash
Use it to find the angle, not to write the whole memo.
What you can hand it, and what you’ll still have to catchClaude Sonnet 5.5
Makes the right launch call. Double-check its facts.
What you can hand it, and what you’ll still have to catch