Skip to content

Published by our team as UserApproved.ai · August 2026

Better decisions
before you experiment

UserApproved outscored Claude and GPT on every quality dimension in a blind review by three independent CRO experts

See the results

One ecommerce store · One run per system

Overall quality Score out of 100
85/100

The highest-rated
growth analysis in this run

UserApproved85
Purpose-built agents
GPT Sol 5.668
High
Claude Opus 565.2
X-high

Three independent reviewers · Scores out of 100

Fewer findings
More that matter

Three of UserApproved’s five findings were rated critical. Claude and GPT produced twelve findings combined, with one rated critical.

UserApproved60% critical

3 of 5 findings rated critical

GPT Sol 5.60% critical

0 of 5 findings rated critical

Claude Opus 514% critical

1 of 7 findings rated critical

Each mark represents one finding Rated critical

What makes a finding
worth acting on

Each finding was rated on a four-point scale across evidence, reasoning, usefulness and the next action

UserApprovedGPT Sol 5.6Claude Opus 5Reviewer scores · 1–4

Evidence your team can act on

Trace the finding back to the behavior, artifact or number behind it

Evidence grounded
UserApproved3.20
GPT2.57
Claude2.23

The why behind the behavior

Connect observations into an explanation that shapes the next experiment

Strategic synthesis
UserApproved3.50
GPT2.93
Claude2.85

Work worth prioritizing

Focus team attention on findings that can change a growth decision

Impact and usefulness
UserApproved3.20
GPT2.63
Claude2.35

A clear next experiment

Know what to change, where to change it and what to measure

Action specificity
UserApproved3.70
GPT2.76
Claude3.00

Same brief
Independent judgment

On August 4, 2026, all three systems investigated the same live US ecommerce store with the same available data, APIs, MCP integrations and browser tools. Three external CRO professionals scored the outputs without knowing which system produced them.

The brief

Act as an ecommerce growth expert and identify actionable revenue opportunities

Scoring, coverage and cost

Each finding was scored from 1 to 4 on each dimension, for a maximum of 16 points, then normalized to 100. Published scores are the means across the three reviewers.

4
One minor flaw a reader would not act on differently
3
Passes, with a real and citable flaw
2
Several real flaws, or one that materially weakens the finding
1
Barely passes; a reader would distrust the pillar
Coverage, elapsed time and reported quality per dollar
SystemCoverageElapsedQuality / $
UserApprovedDesktop + mobile · 4 persona paths46:523.3
GPT Sol 5.6 · HighDesktop · 1 primary path40:003.7
Claude Opus 5 · X-highDesktop · 1 primary path39:002.0

GPT had the highest reported quality-per-dollar index. UserApproved had the highest quality score, with broader coverage and the longest elapsed time. Provider accounting can differ, so the cost index applies to this run. The store is available to prospects under NDA.

This benchmark measures CRO analysis quality in one ecommerce run. It does not measure patient enrollment lift or live experiment outcomes.

Put better intelligence
behind your growth

See how we bring patient understanding and experimentation together to turn more interest into enrollments

See the Eden story

Free setup and integration