Comparative study
Sierra: policy-resolution study
Live storefront evaluation of policy-compliant resolution, answer quality and response speed.
Key findings
- Shopping composite, out of 100: Sierra 59.
- Support composite, out of 100: Sierra 61.6.
- Overall composite, the mean of both lanes: Sierra 60.3.
- Average time to a full answer: shopping Sierra 8.5 s; support Sierra 8.4 s.
- Scores cover the AI agent in each store's on-site chat widget across the 5 storefront deployments captured Sep 22–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab.
Results
Scores out of 100 for the storefronts in this study. Each bar splits a lane composite into the points resolution, quality and speed add.
ShoppingComposite / 100
- Sierra8.5 s to a full answer59
SupportComposite / 100
- Sierra8.4 s to a full answer61.6
- Resolution
- Quality
- Speed
Score breakdown
| Measure | Shopping | Support |
|---|---|---|
| Sierra | Sierra | |
| Composite | 59 | 61.6 |
| Resolution | 61.1 | 40 |
| Quality | 48 | 86 |
| Speed score | 71.1 | 71.6 |
| Average full answer | 8.5 s | 8.4 s |
| Storefronts included | 5 of 5 | 5 of 5 |
| Conversations included | 25 of 25 | 25 of 25 |
Full evidence coverage
| Company | Checkpoints: planned / attempted / observed | Recorded submitted / assessed / unassessable | Attained / unverified | Contexts: included / excluded | Stores: included / excluded | Quality eligible: contexts / stores | Captures: original / repaired |
|---|---|---|---|---|---|---|---|
| Sierra | 250 / 183 / 183 | 183 / 183 / 0 | 104 / 38 | 25 / 0 | 5 / 0 | 20 / 4 | 24 / 1 |
| Company | Checkpoints: planned / attempted / observed | Recorded submitted / assessed / unassessable | Attained / unverified | Contexts: included / excluded | Stores: included / excluded | Quality eligible: contexts / stores | Captures: original / repaired |
|---|---|---|---|---|---|---|---|
| Sierra | 250 / 242 / 242 | 242 / 242 / 0 | 96 / 137 | 25 / 0 | 5 / 0 | 24 / 5 | 25 / 0 |
- Policy-compliant resolution · public-session scope
- Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
- Quality
- The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
- Full-answer speed
- Average time to the complete answer, scored 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
- Composite
- Shopping: 40% policy-compliant resolution, 35% quality, 25% speed. Support: 50% policy-compliant resolution, 40% quality, 10% speed.
- Overall composite
- Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.
This study covers the chat widget only
Studies test one form factor: the AI agent in each store’s on-site chat widget. Other products the vendors sell are outside its scope.
Read this before the scores
Policy-compliant resolution can include a verified policy-required next step. It does not establish actual refunds, completed account actions or human-resolved orders. This is a separately versioned Alhena Research Lab method, not an independent certification or the Gorgias leaderboard.
- Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.
- This evaluation used the previously published method, frozen before capture and judging.
- The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.
- Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.
- Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider.
- 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators.
How the scores were checked
Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible.
Read the versioned methodologyWhat changed from the source study
- Replaces the source automation classifier with policy-compliant resolution.
- Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence.
- New automated captures use the frozen protocol; historical study scores remain unchanged.
Audit notes (2)
- Provider names were masked where practicable; full anonymity is not claimed.
- Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not.
Frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b
Read every conversation behind these scores
Detailed decisions, source records and available study downloads require a verified work email. Your email and view or download activity are shared privately with Alhena. Privacy details.
Considering Alhena? See how its shopping and support agents would handle your customers’ questions.
Book a demo