All evaluated tools

TOOL EVALUATION · QUALITY PILOT

Alhena

Shopping and support quality across three selected customer storefronts, using the same 26 published criteria.

alhena.ai
Within 30 days

SHOPPING QUALITY

96/ 100

Mean of three storefront conversations.

SUPPORT QUALITY

100/ 100

Mean of three storefront conversations.

3Storefronts
6Conversations
60Captured turns
26Published criteria

Original capture dates: Sep 20, 2026. Eligible for new comparisons until Oct 20, 2026.

Explore detailed evidence View evaluation availability

Summaries are public. Detailed conversations and scoring decisions require a verified work email.

BEHIND THE TOOL SCORE

Three deployments. Every result visible.

Customer storefrontShopping /100Support /100Captured
Tatcha98100Sep 20, 2026
Paula's Choice95100Sep 20, 2026
Benchmade95100Sep 20, 2026

SIDE BY SIDE

Alhena comparison reports

Reports retain their own capture dates and source evaluations. Historical reports may use an earlier evaluation than this profile.

QUALITY COMPARISONPublished Sep 20, 2026

Alhena vs. Gorgias

Shopping and support quality across six live storefronts, scored against Gorgias’s published criteria.

6 storefronts120 captured turns26 criteria
Explore the report

Original capture dates and limitations are retained. Read the evidence before interpreting the scores.

QUALITY SCORES / 100

Shopping

Alhena96
Gorgias76.7

Support

Alhena100
Gorgias82.7
Scope and limitations

Operated and commissioned by Alhena Research Lab. Scores describe this selected sample, not universal tool performance or an overall vendor ranking.

Small selected sample: three stores per provider and one conversation per lane/store; not a randomized or matched-store study.

Different stores, catalogs, configurations, browser sessions and current dates limit causal vendor comparisons.

Only two fixed themes were run, not all five per lane or the separate guardrail battery.

Observed response completion substitutes for timed-answer eligibility for quality-only grading. No reliable live latency, automation/speed composite, statistical superiority or official rank is claimed.

Three conversations per vendor/lane are below the published 15-judged-conversation threshold.

Alhena lanes used New Chat in a shared profile. Paula’s Choice claimed preferences not supplied by the script; the record cannot distinguish retained context from unsupported assumptions. Cold isolation was not established.

Store/vendor masking was limited; URLs and product names could reveal identity.

Primary and audit passes were AI agents, not external independent researchers. Audit used full judging packets rather than upstream 450-character tails.

Policy and link checks are separately disclosed; no extra scoring criteria or penalties were introduced. A 100 quality score does not certify factual perfection.

Historical current-rubric regrading cannot reconstruct an unrecorded historical model/prompt version and does not prove manipulation.

Some conversations are reused unchanged from published reports. Their original capture dates and source limitations remain applicable; they were not recaptured or rejudged for this comparison.

Historical reused captures inherit the original published source attribution. Per-turn AI-author proof was not recorded by that source and has not been invented or newly verified.

Read the scoring rubric and methodology