# Alhena Research Lab > Published evaluations of ecommerce AI shopping and support agents on live storefronts, operated by Alhena. Every tool gets the same customer conversations; scores cover policy-compliant resolution, answer quality and full-answer speed, with conversation evidence behind each score. Everything below in one file, including every study's scores, limitations and the FAQ: https://evals.alhena.ai/llms-full.txt ## Latest result per tool - In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Alhena scored 68.4 for shopping and 81.4 for support, as composite scores out of 100. Full answers arrived in 11.8 s for shopping and 8.7 s for support on average. - In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Gorgias scored 55.8 for shopping and 75.8 for support, as composite scores out of 100. Full answers arrived in 17.2 s for shopping and 14.1 s for support on average. - In Sierra: policy-resolution study (captured Sep 22–23, 2026), Sierra scored 59 for shopping and 61.6 for support, as composite scores out of 100. Full answers arrived in 8.5 s for shopping and 8.4 s for support on average. - Scores cover the AI agent in each store's on-site chat widget on the storefronts in each tool's study, not other products the vendors sell, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. ## Current scoring: policy-resolution-v1 The homepage and current tool profiles use the newest published compatible policy-resolution study for each provider, selected by capture date. Shopping and support COMPOSITE scores combine policy-compliant resolution, answer quality and full-answer speed. Quality is a distinct component and must not be used as a label for a composite. Shopping weights: resolution 40%, quality 35%, speed 25%. Support weights: resolution 50%, quality 40%, speed 10%. Overall is the equal mean of the unrounded lane composites when both are eligible. Resolution measures correct answers or verified merchant-prescribed next steps in public sessions. It does not establish completed refunds, account actions or human-resolved cases. This is Alhena Research Lab's method, not Gorgias's automation ranking or independent third-party certification. Registered and included samples, exclusions, capture dates and limitations must accompany comparisons. Profiles may draw from different studies; use each linked study's sample and dates. ## Public sources - [Current results and historical archive](https://evals.alhena.ai/): latest scores per tool, comparative studies and the quality-pilot archive - [Scoring methodology](https://evals.alhena.ai/methodology): what resolution, quality and speed measure and how composites are weighted - [Complete current method](https://evals.alhena.ai/studies/policy-resolution-v1): the versioned policy-resolution-v1 protocol - [Current machine-readable tool scores, explicit metrics, v2](https://evals.alhena.ai/tool-scores.json) - [All approved study summaries](https://evals.alhena.ai/study-scores.json): JSON with every published score, coverage figure and limitation - [Study catalog](https://evals.alhena.ai/studies) - [Full reference for language models](https://evals.alhena.ai/llms-full.txt) - [Sitemap](https://evals.alhena.ai/sitemap.xml) ## Next steps - [Analyze your tool](https://evals.alhena.ai/request): request an evaluation by naming a tool and three stores that use it - [Book a demo of Alhena](https://alhena.ai/schedule-demo?utm_source=evals.alhena.ai&utm_medium=referral&utm_campaign=research_lab&utm_content=llms): see Alhena's shopping and support agents on your store ## Current tool profiles - [Alhena](https://evals.alhena.ai/tools/alhena-ai-e8955b5e31) — policy-resolution-v1; captures 2026-09-21T14:13:43.224Z to 2026-09-22T00:30:54.972Z; source [Alhena and Gorgias: policy-compliant resolution](https://evals.alhena.ai/studies/alhena-gorgias-policy-resolution-2026-09-21) - [Gorgias](https://evals.alhena.ai/tools/gorgias-com-b56a925583) — policy-resolution-v1; captures 2026-09-21T14:13:43.224Z to 2026-09-22T00:30:54.972Z; source [Alhena and Gorgias: policy-compliant resolution](https://evals.alhena.ai/studies/alhena-gorgias-policy-resolution-2026-09-21) - [Sierra](https://evals.alhena.ai/tools/sierra-ai-98c6f6e402) — policy-resolution-v1; captures 2026-09-22T15:03:08.664Z to 2026-09-23T00:53:06.466Z; source [Sierra: policy-resolution study](https://evals.alhena.ai/studies/sierra-ai-98c6f6e402-9cba991e) ## Published studies - [Alhena vs. Sierra: policy-resolution comparison](https://evals.alhena.ai/studies/alhena-ai-e8955b5e31-vs-sierra-ai-98c6f6e402-549df402bc) — published 2026-09-23T00:59:21.323Z; original captures 2026-09-21T14:20:25.398Z to 2026-09-23T00:53:06.466Z - [Gorgias vs. Sierra: policy-resolution comparison](https://evals.alhena.ai/studies/gorgias-com-b56a925583-vs-sierra-ai-98c6f6e402-633b8fd26e) — published 2026-09-23T00:59:21.323Z; original captures 2026-09-21T14:13:43.224Z to 2026-09-23T00:53:06.466Z - [Sierra: policy-resolution study](https://evals.alhena.ai/studies/sierra-ai-98c6f6e402-9cba991e) — published 2026-09-23T00:59:21.323Z; original captures 2026-09-22T15:03:08.664Z to 2026-09-23T00:53:06.466Z - [Alhena and Gorgias: policy-compliant resolution](https://evals.alhena.ai/studies/alhena-gorgias-policy-resolution-2026-09-21) — published 2026-09-22T04:07:09.170Z; original captures 2026-09-21T14:13:43.224Z to 2026-09-22T00:30:54.972Z ## Historical quality pilot and submissions quality-pilot-v1 measures quality only using the 26 pinned criteria: three storefronts, six conversations and two fixed question themes per tool. Its original scores are preserved in the historical archive and are not the current study scores. New Analyze your tool submissions use policy-resolution-v1: the user supplies three customer storefronts, research verifies two more, and operator approval authorizes the five-storefront evaluation and automatic publication after validation. Existing pilot jobs retain their frozen scope. Compatible pilot evidence can be reused within 30 days of its original capture; publication does not refresh that clock. - [Historical quality archive](https://evals.alhena.ai/#archive) - [Original quality-pilot method and criteria](https://evals.alhena.ai/methodology/quality-pilot-v1) - [Historical quality-pilot scores, v1](https://evals.alhena.ai/quality-pilot-scores.json) - [Alhena vs. Gorgias — quality pilot](https://evals.alhena.ai/reports/alhena-vs-gorgias-2026-09-20) Historical quality-only profiles without a published current study: None. ## Evidence access Public summaries, scores and methodology are open. Detailed conversations, criterion decisions, source records and downloadable evidence require a verified work email. No requester identity is included in public datasets. Always cite the original capture dates and sample limitations. Republished or reused evidence is not a fresh evaluation.