Which model is best at which PM task?
The only LLM benchmark that matters to product execs. AI models tested and evaluated on real product work across a broad set of tasks.
7 models13 core tasks182 outputs graded13 of 13 tasks calibrated
| Model | Discover | Design | Define | Experiment | Challenge | Operate | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 SolOpenAI | 89#1 | 94Best | 93 | 89 | 100Best | 87 | 87 | 69 | 88Best | 93 | 88 | 98Best | 80 | 88Best |
| GPT-6 AstraOpenAI | 88#2 | 86 | 95Best | 70 | 98 | 88 | 81 | 95Best | 73 | 98Best | 96Best | 95 | 88Best | 88 |
| GPT-6 LunaOpenAI | 81#3 | 61 | 84 | 85 | 90 | 77 | 89Best | 67 | 78 | 92 | 82 | 93 | 79 | 79 |
| Sonnet 5.5Anthropic | 80#4 | 86 | 78 | 93Best | 58 | 95Best | 84 | 76 | 84 | 80 | 74 | 81 | 63 | 87 |
| Opus 5.5Anthropic | 72#5 | 74 | 69 | 77 | 71 | 72 | 74 | 64 | 80 | 62 | 81 | 64 | 64 | 84 |
| Gemini 3.8 FlashGoogle | 52#6 | 73 | 51 | 76 | 56 | 58 | 52 | 38 | 47 | 37 | 48 | 48 | 41 | 47 |
| Gemini 3.5 Flash-LiteGoogle | 45#7 | 38 | 41 | 49 | 59 | 17 | 58 | 43 | 25 | 53 | 37 | 69 | 50 | 44 |
Inside the benchmark
The short version of every page. Each one goes a lot deeper.
The check AI fails most: in customer research call guide, “marks what to cut if the call runs over” passes just 29% of the time.
What AI gets right and wrong, task by taskThe most common mistake: hypothesis stated as fact, in 26 reviewed outputs.
- 1Hypothesis stated as factReframe it as a hypothesis26
- 2Invented evidenceVerify or remove the claim22
- 3Numbers wrongRedo the arithmetic13
- 4Test or gate too weakTighten the test13
All 13 tasks are calibrated: the graders agree with our PM.
- Customer research call guide
- Extract discovery insights
- One-shot prototype
- Activation & onboarding review
- Build a roadmap
- Develop product strategy
- Write a PRD
- Experiment specification
- Analyse experiment results
- 1000x an idea
- Challenge an idea
- Write a stakeholder update
- Make the launch call
Same briefTwo graders, one checklistA PM checks the checkers
How the scores workRecent drops
Gemini 3.8 Flash
Use it to find the angle, not to write the whole memo.
Gemini 3.8 Flash is good at spotting what a brief is really about. It found the mechanism hidden in the data on both 1000x cases, built roadmaps around their dependencies and a hard deadline, and wrote a customer call guide you could use with a quick edit. The trouble is everything around those insights: it fills gaps in the brief with facts of its own, and states its guesses as findings. That's why the LLM judge found only 1 of its 24 graded written outputs usable with a quick edit, and five of its 26 outputs made a critical mistake.
| Model | OverallAll core tasks | Discover | Design | Define | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Customer research call guide | Extract discovery insights | One-shot prototype | Activation & onboarding review | Build a roadmap | Develop product strategy | Write a PRD | Experiment specification | Analyse experiment results | 1000x an idea | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6.1 SolOpenAI | 89#1 | 94Best | 93 | 89 | 100Best | 87 | 87 | 69 | 88Best | 93 | 88 | 98Best | 80 | 88Best |
| Gemini 3.8 FlashGoogle | 52#6 | 73 | 51 | 76 | 56 | 58 | 52 | 38 | 47 | 37 | 48 | 48 | 41 | 47 |
- Best setupGemini 3.8 Flash · API51.7#6 overall
- Against the leader−37.0GPT-6.1 Sol · API
Claude Sonnet 5.5
Makes the right launch call. Double-check its facts.
Sonnet 5.5 is at its best when the job is a judgement call. It made the strongest launch calls of any model we've tested, and its challenges go straight for the assumption that matters. The catch is how it gets there: it often states an interpretation as fact, or uses a figure the brief never gave it. That's why the LLM judge found only 7 of its 13 graded outputs usable with a quick edit. Budget time to check its evidence.
Written 30 Sept 2026. We’ve added 4 tasks since (Build a roadmap, Customer research call guide, Experiment specification and 1000x an idea), so the scores here cover more than the write-up does.
| Model | OverallAll core tasks | Discover | Design | Define | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Customer research call guide | Extract discovery insights | One-shot prototype | Activation & onboarding review | Build a roadmap | Develop product strategy | Write a PRD | Experiment specification | Analyse experiment results | 1000x an idea | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6.1 SolOpenAI | 89#1 | 94Best | 93 | 89 | 100Best | 87 | 87 | 69 | 88Best | 93 | 88 | 98Best | 80 | 88Best |
| Sonnet 5.5Anthropic | 80#4 | 86 | 78 | 93Best | 58 | 95Best | 84 | 76 | 84 | 80 | 74 | 81 | 63 | 87 |
- Best setupSonnet 5.5 · API79.7#4 overall
- Against the leader−8.9GPT-6.1 Sol · API
GPT-6.1 Sol
Great at the thinking. Check the follow-through.
GPT-6.1 Sol is very good at the part of PM work that needs judgement. 15 of its 16 graded outputs are usable with at most a quick edit, and no model we've tested is better at challenging an idea. Where it slips is the follow-through: the rollback trigger on a launch call, the trade-off rule in a PRD, the "here's what we've already done" in an exec update.
Written 30 Sept 2026. We’ve added 4 tasks since (Build a roadmap, Customer research call guide, Experiment specification and 1000x an idea), so the scores here cover more than the write-up does.
| Model | OverallAll core tasks | Discover | Design | Define | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Customer research call guide | Extract discovery insights | One-shot prototype | Activation & onboarding review | Build a roadmap | Develop product strategy | Write a PRD | Experiment specification | Analyse experiment results | 1000x an idea | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6.1 SolOpenAI | 89#1 | 94Best | 93 | 89 | 100Best | 87 | 87 | 69 | 88Best | 93 | 88 | 98Best | 80 | 88Best |
- Best setupGPT-6.1 Sol · API88.7#1 overall
Know what a new model gets wrong before you rely on it
Every few weeks a new model claims to change everything. We put each one through real PM work, grade every output blind, and send you one short email: what you can hand it, and the mistakes you’ll still have to catch. Better to find out here than in your own work.