All studies

Comparative study

Sierra: policy-resolution study

Live storefront evaluation of policy-compliant resolution, answer quality and response speed.

Published . Captured . Commissioned by Alhena Research Lab using the policy-resolution-v1 method.

5storefront deployments
50conversations captured
50conversations judged
5guardrail tests
425checkpoints judged and audited

Key findings

  • Shopping composite, out of 100: Sierra 59.
  • Support composite, out of 100: Sierra 61.6.
  • Overall composite, the mean of both lanes: Sierra 60.3.
  • Average time to a full answer: shopping Sierra 8.5 s; support Sierra 8.4 s.
  • Scores cover the AI agent in each store's on-site chat widget across the 5 storefront deployments captured Sep 22–23, 2026. They do not cover other products either vendor sells, every deployment, or an overall vendor ranking. Alhena operates Alhena Research Lab.

Results

Scores out of 100 for the storefronts in this study. Each bar splits a lane composite into the points resolution, quality and speed add.

Sierra60.3overall composite

ShoppingComposite / 100

  • Sierra8.5 s to a full answer
    59

SupportComposite / 100

  • Sierra8.4 s to a full answer
    61.6
  • Resolution
  • Quality
  • Speed

Score breakdown

Each measure out of 100 unless marked, by lane and provider.
MeasureShoppingSupport
SierraSierra
Composite5961.6
Resolution61.140
Quality4886
Speed score71.171.6
Average full answer8.5 s8.4 s
Storefronts included5 of 55 of 5
Conversations included25 of 2525 of 25
Full evidence coverage
Shopping evidence coverage
CompanyCheckpoints: planned / attempted / observedRecorded submitted / assessed / unassessableAttained / unverifiedContexts: included / excludedStores: included / excludedQuality eligible: contexts / storesCaptures: original / repaired
Sierra250 / 183 / 183183 / 183 / 0104 / 3825 / 05 / 020 / 424 / 1
Support evidence coverage
CompanyCheckpoints: planned / attempted / observedRecorded submitted / assessed / unassessableAttained / unverifiedContexts: included / excludedStores: included / excludedQuality eligible: contexts / storesCaptures: original / repaired
Sierra250 / 242 / 242242 / 242 / 096 / 13725 / 05 / 024 / 525 / 0
Policy-compliant resolution · public-session scope
Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
Quality
The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
Full-answer speed
Average time to the complete answer, scored 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
Composite
Shopping: 40% policy-compliant resolution, 35% quality, 25% speed. Support: 50% policy-compliant resolution, 40% quality, 10% speed.
Overall composite
Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.

This study covers the chat widget only

Studies test one form factor: the AI agent in each store’s on-site chat widget. Other products the vendors sell are outside its scope.

Read this before the scores

Policy-compliant resolution can include a verified policy-required next step. It does not establish actual refunds, completed account actions or human-resolved orders. This is a separately versioned Alhena Research Lab method, not an independent certification or the Gorgias leaderboard.

  • Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.
  • This evaluation used the previously published method, frozen before capture and judging.
  • The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.
  • Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.
  • Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider.
  • 1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators.

How the scores were checked

Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible.

Read the versioned methodology

What changed from the source study

  • Replaces the source automation classifier with policy-compliant resolution.
  • Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence.
  • New automated captures use the frozen protocol; historical study scores remain unchanged.
Audit notes (2)
  • Provider names were masked where practicable; full anonymity is not claimed.
  • Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not.

Frozen method SHA-256 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b

Read every conversation behind these scores

Detailed decisions, source records and available study downloads require a verified work email. Your email and view or download activity are shared privately with Alhena. Privacy details.

Considering Alhena? See how its shopping and support agents would handle your customers’ questions.

Book a demo