ALHENA RESEARCH LAB

Put ecommerce AI
to the test.

Explore every evaluated tool. Compare shopping and support quality, then follow the scores to real conversation evidence.

One tool. Three customer storefronts. A place in the research library.

THE GROWING EVIDENCE LIBRARY

Evaluate once.
Compare across the field.

2evaluated tools
1comparison reports
26quality criteria

Every comparison uses validated source evaluations. Original capture dates and limitations travel with the evidence.

Shopping qualitySupport quality
Alhena96 / 100100 / 100Gorgias76.7 / 10082.7 / 100
Real storefrontsThree deployments per tool
One published rubricFixed questions and weights
Two quality scoresShopping and support, separately

THE RESULTS, OPEN TO EXPLORE

See the tools. Compare the evidence.

Scores describe the selected storefront sample. They are not an overall vendor ranking.

Latest complete evaluation for each tool. Scores out of 100.
ToolShopping qualitySupport qualitySampleCapture datesFreshness
Alhenaalhena.ai
96
100
3 storefronts
6 conversations
Sep 20, 2026Within 30 days
Gorgiaswww.gorgias.com
76.7
82.7
3 storefronts
6 conversations
Sep 20, 2026Within 30 days
Same criteria. Visible limits.

Each tool has six ten-turn conversations across three customer storefronts: three shopping and three support conversations. Charts use one complete evaluation per tool, so reusing it in multiple reports does not inflate the sample.

Fresh means every capture is within 30 days. Older results stay visible with a refresh notice, but are not used to create new automatic comparisons. Different storefronts and merchant configurations can affect scores.

Read all 26 scoring criteria

ADD TO THE EVIDENCE

How does your tool perform?

Enter one tool and three customer storefronts. After review, we evaluate it once and create comparisons against compatible, recent evaluations already in the library.

Analyze your tool

A FEW FAIR QUESTIONS

Know what you’re looking at.

Do new comparisons require new testing?

The new tool is tested only on missing or expired storefront conversations. Pairwise reports reuse validated source evaluations captured within 30 days. Assembling a comparison adds no new shopper conversations or model judging calls.

What can I read without signing in?

Tool scores, sample sizes, capture dates, summaries and the rubric are public. Detailed conversations, criterion decisions and evidence downloads require a verified work email.

Who runs Alhena Research Lab?

Alhena operates and commissions these studies. Separate AI judging and audit do not make the Lab an independent research institution. The scoring record and limitations accompany every report.

Are these the full Gorgias leaderboard scores?

No. These studies apply the pinned shopping and support quality criteria with two fixed themes. They do not calculate automation, speed or an overall composite score.