Alhena vs. Gorgias
Shopping and support quality across six live storefronts, scored against Gorgias’s published criteria.
Original capture dates and limitations are retained. Read the evidence before interpreting the scores.
QUALITY SCORES / 100
TOOL EVALUATION · QUALITY PILOT
Shopping and support quality across three selected customer storefronts, using the same 26 published criteria.
www.gorgias.comSHOPPING QUALITY
Mean of three storefront conversations.
SUPPORT QUALITY
Mean of three storefront conversations.
Original capture dates: Sep 20, 2026. Eligible for new comparisons until Oct 20, 2026.
Summaries are public. Detailed conversations and scoring decisions require a verified work email.
BEHIND THE TOOL SCORE
| Customer storefront | Shopping /100 | Support /100 | Captured |
|---|---|---|---|
| Beekman 1802 | 73 | 68 | Sep 20, 2026 |
| Shoebacca | 73 | 100 | Sep 20, 2026 |
| Nordic Outdoor | 84 | 80 | Sep 20, 2026 |
SIDE BY SIDE
Reports retain their own capture dates and source evaluations. Historical reports may use an earlier evaluation than this profile.
Shopping and support quality across six live storefronts, scored against Gorgias’s published criteria.
Original capture dates and limitations are retained. Read the evidence before interpreting the scores.
QUALITY SCORES / 100
Operated and commissioned by Alhena Research Lab. Scores describe this selected sample, not universal tool performance or an overall vendor ranking.
Small selected sample: three stores per provider and one conversation per lane/store; not a randomized or matched-store study.
Different stores, catalogs, configurations, browser sessions and current dates limit causal vendor comparisons.
Only two fixed themes were run, not all five per lane or the separate guardrail battery.
Observed response completion substitutes for timed-answer eligibility for quality-only grading. No reliable live latency, automation/speed composite, statistical superiority or official rank is claimed.
Three conversations per vendor/lane are below the published 15-judged-conversation threshold.
Alhena lanes used New Chat in a shared profile. Paula’s Choice claimed preferences not supplied by the script; the record cannot distinguish retained context from unsupported assumptions. Cold isolation was not established.
Store/vendor masking was limited; URLs and product names could reveal identity.
Primary and audit passes were AI agents, not external independent researchers. Audit used full judging packets rather than upstream 450-character tails.
Policy and link checks are separately disclosed; no extra scoring criteria or penalties were introduced. A 100 quality score does not certify factual perfection.
Historical current-rubric regrading cannot reconstruct an unrecorded historical model/prompt version and does not prove manipulation.
Some conversations are reused unchanged from published reports. Their original capture dates and source limitations remain applicable; they were not recaptured or rejudged for this comparison.
Historical reused captures inherit the original published source attribution. Per-turn AI-author proof was not recorded by that source and has not been invented or newly verified.
Read the scoring rubric and methodology