Evidence your team can act on
Trace the finding back to the behavior, artifact or number behind it
Evidence groundedPublished by our team as UserApproved.ai · August 2026
UserApproved outscored Claude and GPT on every quality dimension in a blind review by three independent CRO experts
See the resultsOne ecommerce store · One run per system
The highest-rated
growth analysis in this run
Three independent reviewers · Scores out of 100
Three of UserApproved’s five findings were rated critical. Claude and GPT produced twelve findings combined, with one rated critical.
3 of 5 findings rated critical
0 of 5 findings rated critical
1 of 7 findings rated critical
Each finding was rated on a four-point scale across evidence, reasoning, usefulness and the next action
Trace the finding back to the behavior, artifact or number behind it
Evidence groundedConnect observations into an explanation that shapes the next experiment
Strategic synthesisFocus team attention on findings that can change a growth decision
Impact and usefulnessKnow what to change, where to change it and what to measure
Action specificityOn August 4, 2026, all three systems investigated the same live US ecommerce store with the same available data, APIs, MCP integrations and browser tools. Three external CRO professionals scored the outputs without knowing which system produced them.
Act as an ecommerce growth expert and identify actionable revenue opportunities
Each finding was scored from 1 to 4 on each dimension, for a maximum of 16 points, then normalized to 100. Published scores are the means across the three reviewers.
| System | Coverage | Elapsed | Quality / $ |
|---|---|---|---|
| UserApproved | Desktop + mobile · 4 persona paths | 46:52 | 3.3 |
| GPT Sol 5.6 · High | Desktop · 1 primary path | 40:00 | 3.7 |
| Claude Opus 5 · X-high | Desktop · 1 primary path | 39:00 | 2.0 |
GPT had the highest reported quality-per-dollar index. UserApproved had the highest quality score, with broader coverage and the longest elapsed time. Provider accounting can differ, so the cost index applies to this run. The store is available to prospects under NDA.
This benchmark measures CRO analysis quality in one ecommerce run. It does not measure patient enrollment lift or live experiment outcomes.
See how we bring patient understanding and experimentation together to turn more interest into enrollments
See the Eden storyFree setup and integration