# Alhena Research Lab: full reference > Published evaluations of ecommerce AI shopping and support agents on live storefronts, operated by Alhena. Every tool gets the same customer conversations. Scores cover policy-compliant resolution, answer quality and full-answer speed, with conversation evidence behind each score. Source: https://evals.alhena.ai/. Index: https://evals.alhena.ai/llms.txt. When citing a score, cite the study URL and its capture dates. Scores describe the storefronts in each study, not every deployment or an overall vendor ranking. Alhena operates the Lab, and its own agent appears in the results. ## What every score is made of Shopping and support are scored separately, each out of 100. - Policy-compliant resolution: did the customer get the correct answer, or the next step the merchant's published policy requires, such as a handoff to a person? It does not establish completed refunds, account actions or human-resolved orders. - Answer quality: the 26 published quality criteria (16 for shopping, 10 for support), covering direct, specific answers grounded in the store's catalog and policies, useful product recommendations and clear next steps. - Full-answer speed: time until the complete answer appears. 3 seconds scores 100; 22 seconds or more scores 0. - Composite weights: shopping 40% resolution, 35% quality, 25% speed; support 50% resolution, 40% quality, 10% speed. The overall composite is the mean of the two lane composites when both are eligible. ## How an evaluation works 1. Pick live storefronts: five stores where the tool is deployed. The vendor suggests three; the Lab finds and verifies two more from published customer stories. 2. Talk to the agent like a customer: every store gets the same scripted conversations, five shopping and five support, plus a guardrail test that tries to pull the agent off course. Conversation themes: - Shopping: Everyday value buyer; Gift shopper; Specific need / reassurance; Comparison, budget-tight; Total beginner. - Support: Order tracking / delay; Returns & exchanges; Damaged / faulty item; Modify / cancel order; Shipping & returns policy. Question pools come from Gorgias's public AI agent benchmark (https://github.com/gorgias/ai-agent-benchmark). 3. Judge every checkpoint twice, blind: each checkpoint gets an AI judgment and a separate blind audit, and only counts when both agree on the evidence. Scores come from fixed arithmetic. 4. Publish the evidence: scores, capture dates, coverage and limits are public. Detailed conversations and scoring decisions require a verified work email. Full methodology: https://evals.alhena.ai/methodology and https://evals.alhena.ai/studies/policy-resolution-v1. ## Published studies ### Alhena vs. Sierra: policy-resolution comparison - URL: https://evals.alhena.ai/studies/alhena-ai-e8955b5e31-vs-sierra-ai-98c6f6e402-549df402bc - Published: Sep 23, 2026. Captured: Sep 21–23, 2026 (2026-09-21T14:20:25.398Z to 2026-09-23T00:53:06.466Z). - Method: policy-resolution-v1, frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b. Commissioned by Alhena Research Lab. - Sample: 100 conversations captured, 100 judged, 10 guardrail tests, 917 checkpoints judged and audited. - Overall composite: Sierra 60.3, Alhena 74.9. Comparison assembled from published evaluations; no new conversations were run for this comparison. Key findings: - Shopping composite, out of 100: Sierra 59, Alhena 68.4. - Support composite, out of 100: Sierra 61.6, Alhena 81.4. - Overall composite, the mean of both lanes: Sierra 60.3, Alhena 74.9. - Average time to a full answer: shopping Sierra 8.5 s, Alhena 11.8 s; support Sierra 8.4 s, Alhena 8.7 s. - Scores cover the AI agent in each store's on-site chat widget across the 10 storefront deployments captured Sep 21–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. | Measure | Shopping: Sierra | Shopping: Alhena | Support: Sierra | Support: Alhena | |---|---:|---:|---:|---:| | Composite (out of 100) | 59 | 68.4 | 61.6 | 81.4 | | Policy-compliant resolution | 61.1 | 64.8 | 40 | 72.8 | | Answer quality | 48 | 83 | 86 | 95 | | Full-answer speed score | 71.1 | 53.7 | 71.6 | 70 | | Average time to a full answer | 8.5 s | 11.8 s | 8.4 s | 8.7 s | | Storefronts included | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 | | Conversations included | 25 of 25 | 25 of 25 | 25 of 25 | 25 of 25 | Scope: this study tests one form factor, the AI agent in each store's on-site chat widget. Other products the vendors sell were not tested. Limitations: - Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification. - This evaluation used the previously published method, frozen before capture and judging. - The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified. - Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage. - Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider. - 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators. - This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments. - All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates. - Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained. - Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established. - The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed. - Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores. - Derived from separate published runs under policy-resolution-v1. Original capture dates and coverage are retained; execution environments may differ. Audit: This comparison retains the original judgments and audit coverage of each source study. No new judgment or audit is claimed for this derived comparison. - Provider names were masked where practicable; full anonymity is not claimed. - Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not. - Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit. - One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts. - The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance. - A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts. - The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance. - The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage. - Source methods: sierra-ai-98c6f6e402-9cba991e (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b); alhena-gorgias-policy-resolution-2026-09-21 (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b). Read each source method before interpreting differences. Changes from the source study: - Replaces the source automation classifier with policy-compliant resolution. - Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence. - New automated captures use the frozen protocol; historical study scores remain unchanged. ### Gorgias vs. Sierra: policy-resolution comparison - URL: https://evals.alhena.ai/studies/gorgias-com-b56a925583-vs-sierra-ai-98c6f6e402-633b8fd26e - Published: Sep 23, 2026. Captured: Sep 21–23, 2026 (2026-09-21T14:13:43.224Z to 2026-09-23T00:53:06.466Z). - Method: policy-resolution-v1, frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b. Commissioned by Alhena Research Lab. - Sample: 100 conversations captured, 88 judged, 10 guardrail tests, 769 checkpoints judged and audited. - Overall composite: Sierra 60.3, Gorgias 65.8. Comparison assembled from published evaluations; no new conversations were run for this comparison. Key findings: - Shopping composite, out of 100: Sierra 59, Gorgias 55.8. - Support composite, out of 100: Sierra 61.6, Gorgias 75.8. - Overall composite, the mean of both lanes: Sierra 60.3, Gorgias 65.8. - Average time to a full answer: shopping Sierra 8.5 s, Gorgias 17.2 s; support Sierra 8.4 s, Gorgias 14.1 s. - Gorgias's results include 4 of 5 storefronts for shopping and 4 of 5 for support. Storefronts that could not be scored are left out, not counted as zero. - Scores cover the AI agent in each store's on-site chat widget across the 10 storefront deployments captured Sep 21–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. | Measure | Shopping: Sierra | Shopping: Gorgias | Support: Sierra | Support: Gorgias | |---|---:|---:|---:|---:| | Composite (out of 100) | 59 | 55.8 | 61.6 | 75.8 | | Policy-compliant resolution | 61.1 | 60.8 | 40 | 68 | | Answer quality | 48 | 72 | 86 | 94 | | Full-answer speed score | 71.1 | 25.3 | 71.6 | 41.6 | | Average time to a full answer | 8.5 s | 17.2 s | 8.4 s | 14.1 s | | Storefronts included | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | | Conversations included | 25 of 25 | 20 of 25 | 25 of 25 | 18 of 25 | Scope: this study tests one form factor, the AI agent in each store's on-site chat widget. Other products the vendors sell were not tested. Limitations: - Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification. - This evaluation used the previously published method, frozen before capture and judging. - The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified. - Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage. - Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider. - 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators. - This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments. - All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates. - Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained. - Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established. - The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed. - Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores. - Derived from separate published runs under policy-resolution-v1. Original capture dates and coverage are retained; execution environments may differ. Audit: This comparison retains the original judgments and audit coverage of each source study. No new judgment or audit is claimed for this derived comparison. - Provider names were masked where practicable; full anonymity is not claimed. - Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not. - Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit. - One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts. - The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance. - A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts. - The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance. - The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage. - Source methods: sierra-ai-98c6f6e402-9cba991e (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b); alhena-gorgias-policy-resolution-2026-09-21 (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b). Read each source method before interpreting differences. Changes from the source study: - Replaces the source automation classifier with policy-compliant resolution. - Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence. - New automated captures use the frozen protocol; historical study scores remain unchanged. ### Sierra: policy-resolution study - URL: https://evals.alhena.ai/studies/sierra-ai-98c6f6e402-9cba991e - Published: Sep 23, 2026. Captured: Sep 22–23, 2026 (2026-09-22T15:03:08.664Z to 2026-09-23T00:53:06.466Z). - Method: policy-resolution-v1, frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b. Commissioned by Alhena Research Lab. - Sample: 50 conversations captured, 50 judged, 5 guardrail tests, 425 checkpoints judged and audited. - Overall composite: Sierra 60.3. Live storefront evaluation of policy-compliant resolution, answer quality and response speed. Key findings: - Shopping composite, out of 100: Sierra 59. - Support composite, out of 100: Sierra 61.6. - Overall composite, the mean of both lanes: Sierra 60.3. - Average time to a full answer: shopping Sierra 8.5 s; support Sierra 8.4 s. - Scores cover the AI agent in each store's on-site chat widget across the 5 storefront deployments captured Sep 22–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. | Measure | Shopping: Sierra | Support: Sierra | |---|---:|---:| | Composite (out of 100) | 59 | 61.6 | | Policy-compliant resolution | 61.1 | 40 | | Answer quality | 48 | 86 | | Full-answer speed score | 71.1 | 71.6 | | Average time to a full answer | 8.5 s | 8.4 s | | Storefronts included | 5 of 5 | 5 of 5 | | Conversations included | 25 of 25 | 25 of 25 | Scope: this study tests one form factor, the AI agent in each store's on-site chat widget. Other products the vendors sell were not tested. Limitations: - Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification. - This evaluation used the previously published method, frozen before capture and judging. - The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified. - Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage. - Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider. - 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators. Audit: Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible. - Provider names were masked where practicable; full anonymity is not claimed. - Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not. Changes from the source study: - Replaces the source automation classifier with policy-compliant resolution. - Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence. - New automated captures use the frozen protocol; historical study scores remain unchanged. ### Alhena and Gorgias: policy-compliant resolution - URL: https://evals.alhena.ai/studies/alhena-gorgias-policy-resolution-2026-09-21 - Published: Sep 22, 2026. Captured: Sep 21–22, 2026 (2026-09-21T14:13:43.224Z to 2026-09-22T00:30:54.972Z). - Method: policy-resolution-v1, frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b. Commissioned by Alhena Research Lab. - Sample: 100 conversations captured, 88 judged, 10 guardrail tests, 836 checkpoints judged and audited. - Overall composite: Alhena 74.9, Gorgias 65.8. A study of shopping and support on ten storefront deployments, with conversation evidence, public merchant policies, fixed quality criteria and separate response timing. Key findings: - Shopping composite, out of 100: Alhena 68.4, Gorgias 55.8. - Support composite, out of 100: Alhena 81.4, Gorgias 75.8. - Overall composite, the mean of both lanes: Alhena 74.9, Gorgias 65.8. - Average time to a full answer: shopping Alhena 11.8 s, Gorgias 17.2 s; support Alhena 8.7 s, Gorgias 14.1 s. - Gorgias's results include 4 of 5 storefronts for shopping and 4 of 5 for support. Storefronts that could not be scored are left out, not counted as zero. - Scores cover the AI agent in each store's on-site chat widget across the 10 storefront deployments captured Sep 21–22, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. | Measure | Shopping: Alhena | Shopping: Gorgias | Support: Alhena | Support: Gorgias | |---|---:|---:|---:|---:| | Composite (out of 100) | 68.4 | 55.8 | 81.4 | 75.8 | | Policy-compliant resolution | 64.8 | 60.8 | 72.8 | 68 | | Answer quality | 83 | 72 | 95 | 94 | | Full-answer speed score | 53.7 | 25.3 | 70 | 41.6 | | Average time to a full answer | 11.8 s | 17.2 s | 8.7 s | 14.1 s | | Storefronts included | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | | Conversations included | 25 of 25 | 20 of 25 | 25 of 25 | 18 of 25 | Scope: this study tests one form factor, the AI agent in each store's on-site chat widget. Other products the vendors sell were not tested. Product scope beyond the chat widget, from each vendor's public website as of 2026-09-22 (reference information, not study results): | Capability | Alhena | Gorgias | In this study | |---|---|---|---| | AI shopping assistant in the on-site chat widget | yes (AI Shopping Assistant) | yes (AI Shopping Assistant) | tested | | AI support agent in the on-site chat widget | yes (AI Support Concierge) | yes (AI Agent) | tested | | AI visibility in answer engines (AEO/GEO) | yes (AI Visibility) | not listed on gorgias.com | not tested | | AI-written FAQs on product pages | yes (Smart FAQs and the AEO FAQ Engine) | not listed on gorgias.com | not tested | | Conversational product search | yes (Conversational Search) | yes (Search to revenue, in the AI Shopping Assistant) | only inside chat conversations | | Guided and visual shopping (shade matching, similar products, virtual try-on) | yes (Guided Discovery, Shade Matcher, Similar Product Finder, virtual try-ons) | not listed on gorgias.com | not tested | | Proactive nudges and offers | yes (Conversion Nudges) | yes (Convert) | not tested | | Helpdesk and ticketing for human agents | via integrations (Zendesk, Gorgias, Salesforce Service Cloud, Freshdesk and more) | yes (Helpdesk) | not tested | | Help center | not listed on alhena.ai | yes (Help Center) | not tested | | Email support | yes (Email integration) | yes (Helpdesk) | not tested | | SMS | via integrations (through a connected helpdesk) | yes (SMS) | not tested | | Voice | yes (Voice AI) | yes (Voice) | not tested | | WhatsApp | yes (AI Social Commerce) | yes (WhatsApp) | not tested | | Instagram and Facebook | yes (AI Social Commerce) | yes (Social media) | not tested | | Commerce platforms listed | yes (Shopify, WooCommerce, Salesforce Commerce Cloud) | yes (Shopify, BigCommerce, Magento, WooCommerce, PrestaShop) | not tested | Sources: https://alhena.ai/products/ai-shopping-assistant, https://alhena.ai/products/ai-support-concierge, https://alhena.ai/products/ai-visibility, https://alhena.ai/, https://alhena.ai/solutions/industries/beauty-and-skincare, https://alhena.ai/integrations, https://alhena.ai/integrations/email, https://alhena.ai/integrations/gorgias, https://alhena.ai/products/voice-ai, https://alhena.ai/products/ai-social-commerce, https://www.gorgias.com/ai-shopping-assistant, https://www.gorgias.com/ai-agent, https://www.gorgias.com/product/convert, https://www.gorgias.com/product/helpdesk, https://www.gorgias.com/product/help-center, https://www.gorgias.com/product/sms, https://www.gorgias.com/products/voice, https://www.gorgias.com/product/whatsapp, https://www.gorgias.com/product/social-media, https://www.gorgias.com/pricing Alhena products this study doesn't compare: - AI Visibility (AEO/GEO): Measures how your products appear inside AI-generated shopping answers and gives product-level actions to improve visibility. (https://alhena.ai/products/ai-visibility) - Smart FAQs and the AEO FAQ Engine: AI-written answers embedded on every product page, plus citation-ready Q&A pairs built for answer engines. (https://alhena.ai/) - Conversational Search: Natural-language product search for layered queries. The study’s shopping conversations asked product questions but did not score search on its own. (https://alhena.ai/products/ai-shopping-assistant) - Guided and visual shopping: Guided Discovery, Shade Matcher, Similar Product Finder, and virtual try-ons for fit and color. (https://alhena.ai/solutions/industries/beauty-and-skincare) Limitations: - Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification. - This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments. - The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified. - Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage. - All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates. - Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained. - Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established. - The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed. - Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores. Audit: Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible. - Provider names were masked where practicable; full anonymity is not claimed. - Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit. - One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts. - The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance. - A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts. - The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance. - The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage. Changes from the source study: - Replaces the source automation classifier with policy-compliant resolution. - Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence. - Cause-selected capture repairs retained: 12. The source study remains unchanged. ## Current tool profiles - [Alhena](https://evals.alhena.ai/tools/alhena-ai-e8955b5e31): In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Alhena scored 68.4 for shopping and 81.4 for support, as composite scores out of 100. Full answers arrived in 11.8 s for shopping and 8.7 s for support on average. - [Gorgias](https://evals.alhena.ai/tools/gorgias-com-b56a925583): In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Gorgias scored 55.8 for shopping and 75.8 for support, as composite scores out of 100. Full answers arrived in 17.2 s for shopping and 14.1 s for support on average. - [Sierra](https://evals.alhena.ai/tools/sierra-ai-98c6f6e402): In Sierra: policy-resolution study (captured Sep 22–23, 2026), Sierra scored 59 for shopping and 61.6 for support, as composite scores out of 100. Full answers arrived in 8.5 s for shopping and 8.4 s for support on average. ## Frequently asked questions ### What are the latest results? In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Alhena scored 68.4 for shopping and 81.4 for support, as composite scores out of 100. Full answers arrived in 11.8 s for shopping and 8.7 s for support on average. In Alhena and Gorgias: policy-compliant resolution (captured Sep 21–22, 2026), Gorgias scored 55.8 for shopping and 75.8 for support, as composite scores out of 100. Full answers arrived in 17.2 s for shopping and 14.1 s for support on average. In Sierra: policy-resolution study (captured Sep 22–23, 2026), Sierra scored 59 for shopping and 61.6 for support, as composite scores out of 100. Full answers arrived in 8.5 s for shopping and 8.4 s for support on average. Scores cover the AI agent in each store's on-site chat widget on the storefronts in each tool's study, not other products the vendors sell, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab. (https://evals.alhena.ai/studies) ### Does this compare everything Alhena and Gorgias offer? No. Studies test one form factor, the AI agent in each store's on-site chat widget. Alhena lists AI visibility tools for answer engines (AEO/GEO), AI-written FAQs on product pages, and guided and visual shopping such as shade matching and virtual try-on, which gorgias.com does not list. Gorgias lists a helpdesk for human agents, a help center, and SMS, which alhena.ai does not sell as its own products. Alhena connects to a helpdesk for human agents and SMS through integrations instead. None of these were tested. The study's product scope table lists each vendor's offerings with sources. (https://evals.alhena.ai/studies/alhena-gorgias-policy-resolution-2026-09-21#product-scope) ### How is each AI agent tested? Each tool is tested on five live storefronts where it is deployed: the vendor suggests three and the Lab verifies two more. Every store gets the same scripted shopping and support conversations, plus a guardrail test. Each checkpoint gets an AI judgment and a separate blind audit, and scores come from fixed arithmetic. (https://evals.alhena.ai/methodology) ### Is this a ranking of vendors? No. Each score describes the storefronts and capture dates in its study. Merchant setups differ, and storefronts that could not be scored are left out rather than counted as zero. ### What does the composite score measure? It combines resolution, answer quality and full-answer speed with published weights: 40/35/25 for shopping and 50/40/10 for support. Use the metric selector to inspect each part. Resolution can credit a documented next step, such as a required handoff; it does not confirm that a refund or case was completed afterwards. ### Who runs Alhena Research Lab? Alhena operates and commissions these studies, and its own agent appears in the results. Separate AI judging and audits do not make the Lab an independent research institution, so the scoring record and limitations accompany every report. ### What can I read without signing in? Study scores, sample sizes, capture dates, summaries and methods are public. Detailed conversations, criterion decisions and evidence downloads require a verified work email. ### Is every evaluation fully automatic? Research verifies five customer deployments before an operator approves testing. AI then runs the conversations, judges responses and audits resolution decisions, and validated results publish automatically. Temporary failures can retry within fixed limits; unsupported widgets, uncertain response attribution and failed validation may need operator review. Incomplete runs stay unpublished. ### How do I get my AI agent evaluated? Name the tool and three stores that use it, then verify your work email. The Lab researches two more storefronts, runs the evaluation after operator review, and publishes results that pass validation. (https://evals.alhena.ai/request) ### What happened to the earlier quality scores? They remain in the historical quality pilot archive with their original scores and capture dates. The pilot used a separate scope and protocol, so its scores are not combined with the current results. ## Machine-readable data - Current tool scores: https://evals.alhena.ai/tool-scores.json - Published study summaries: https://evals.alhena.ai/study-scores.json - Historical quality-pilot scores: https://evals.alhena.ai/quality-pilot-scores.json - Sitemap: https://evals.alhena.ai/sitemap.xml ## Next steps - Get an AI agent evaluated: https://evals.alhena.ai/request - See Alhena's shopping and support agents on your store: https://alhena.ai/schedule-demo?utm_source=evals.alhena.ai&utm_medium=referral&utm_campaign=research_lab&utm_content=llms_full