Comparative study
Alhena vs. Sierra: policy-resolution comparison
Comparison assembled from published evaluations; no new conversations were run for this comparison.
Key findings
- Shopping composite, out of 100: Sierra 59, Alhena 68.4.
- Support composite, out of 100: Sierra 61.6, Alhena 81.4.
- Overall composite, the mean of both lanes: Sierra 60.3, Alhena 74.9.
- Average time to a full answer: shopping Sierra 8.5 s, Alhena 11.8 s; support Sierra 8.4 s, Alhena 8.7 s.
- Scores cover the AI agent in each store's on-site chat widget across the 10 storefront deployments captured Sep 21–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab.
Results
Scores out of 100 for the storefronts in this study. Each bar splits a lane composite into the points resolution, quality and speed add.
ShoppingComposite / 100
- Sierra8.5 s to a full answer59
- Alhena11.8 s to a full answer68.4
SupportComposite / 100
- Sierra8.4 s to a full answer61.6
- Alhena8.7 s to a full answer81.4
- Resolution
- Quality
- Speed
Score breakdown
| Measure | Shopping | Support | ||
|---|---|---|---|---|
| Sierra | Alhena | Sierra | Alhena | |
| Composite | 59 | 68.4 | 61.6 | 81.4 |
| Resolution | 61.1 | 64.8 | 40 | 72.8 |
| Quality | 48 | 83 | 86 | 95 |
| Speed score | 71.1 | 53.7 | 71.6 | 70 |
| Average full answer | 8.5 s | 11.8 s | 8.4 s | 8.7 s |
| Storefronts included | 5 of 5 | 5 of 5 | 5 of 5 | 5 of 5 |
| Conversations included | 25 of 25 | 25 of 25 | 25 of 25 | 25 of 25 |
Full evidence coverage
| Company | Checkpoints: planned / attempted / observed | Recorded submitted / assessed / unassessable | Attained / unverified | Contexts: included / excluded | Stores: included / excluded | Quality eligible: contexts / stores | Captures: original / repaired |
|---|---|---|---|---|---|---|---|
| Sierra | 250 / 183 / 183 | 183 / 183 / 0 | 104 / 38 | 25 / 0 | 5 / 0 | 20 / 4 | 24 / 1 |
| Alhena | 250 / 242 / 242 | 242 / 242 / 0 | 158 / 64 | 25 / 0 | 5 / 0 | 25 / 5 | 23 / 2 |
| Company | Checkpoints: planned / attempted / observed | Recorded submitted / assessed / unassessable | Attained / unverified | Contexts: included / excluded | Stores: included / excluded | Quality eligible: contexts / stores | Captures: original / repaired |
|---|---|---|---|---|---|---|---|
| Sierra | 250 / 242 / 242 | 242 / 242 / 0 | 96 / 137 | 25 / 0 | 5 / 0 | 24 / 5 | 25 / 0 |
| Alhena | 250 / 250 / 250 | 250 / 250 / 0 | 182 / 63 | 25 / 0 | 5 / 0 | 25 / 5 | 19 / 6 |
- Policy-compliant resolution · public-session scope
- Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
- Quality
- The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
- Full-answer speed
- Average time to the complete answer, scored 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
- Composite
- Shopping: 40% policy-compliant resolution, 35% quality, 25% speed. Support: 50% policy-compliant resolution, 40% quality, 10% speed.
- Overall composite
- Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.
This study covers the chat widget only
Studies test one form factor: the AI agent in each store’s on-site chat widget. Other products the vendors sell are outside its scope.
Read this before the scores
Policy-compliant resolution can include a verified policy-required next step. It does not establish actual refunds, completed account actions or human-resolved orders. This is a separately versioned Alhena Research Lab method, not an independent certification or the Gorgias leaderboard.
- Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.
- This evaluation used the previously published method, frozen before capture and judging.
- The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.
- Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.
- Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider.
- 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators.
- This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments.
- All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates.
- Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained.
- Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established.
- The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed.
- Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores.
- Derived from separate published runs under policy-resolution-v1. Original capture dates and coverage are retained; execution environments may differ.
How the scores were checked
This comparison retains the original judgments and audit coverage of each source study. No new judgment or audit is claimed for this derived comparison.
Read the versioned methodologyWhat changed from the source study
- Replaces the source automation classifier with policy-compliant resolution.
- Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence.
- New automated captures use the frozen protocol; historical study scores remain unchanged.
Audit notes (9)
- Provider names were masked where practicable; full anonymity is not claimed.
- Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not.
- Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit.
- One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts.
- The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance.
- A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts.
- The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance.
- The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage.
- Source methods: sierra-ai-98c6f6e402-9cba991e (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b); alhena-gorgias-policy-resolution-2026-09-21 (712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b). Read each source method before interpreting differences.
Frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b
Read every conversation behind these scores
Detailed decisions, source records and available study downloads require a verified work email. Your email and view or download activity are shared privately with Alhena. Privacy details.
Considering Alhena? See how its shopping and support agents would handle your customers’ questions.
Book a demo