All evaluated tools

LATEST APPROVED STUDY · policy-resolution-v1

Sierra

In Sierra: policy-resolution study (captured Sep 22–23, 2026), Sierra scored 59 for shopping and 61.6 for support, as composite scores out of 100. Full answers arrived in 8.5 s for shopping and 8.4 s for support on average.

sierra.ai

Captures Sep 22, 2026 to Sep 23, 2026. Published Sep 23, 2026. 5 registered storefronts. Selected by capture end date, then publication date.

Public-session scope

Policy-compliant resolution includes a verified policy-required next step. These scores do not prove completed refunds, account actions or human-resolved orders. Coverage differs by lane and provider; this is an Alhena-commissioned study, not independent certification or a universal vendor ranking.

Read the study and evidence Read the methodology

Source: Sierra: policy-resolution study. Summaries are public; detailed evidence requires a verified work email.

Shopping

Composite

59 / 100

40% policy-compliant resolution, 35% quality, 25% speed.

Policy-compliant resolution · public-session scope

61.1 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Answer quality

48 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

71.1 / 100

Full-answer completion 8.5 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

Support

Composite

61.6 / 100

50% policy-compliant resolution, 40% quality, 10% speed.

Policy-compliant resolution · public-session scope

40 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Answer quality

86 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

71.6 / 100

Full-answer completion 8.4 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

Overall composite

Overall composite

60.3 / 100

Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.

Evidence coverage

Registered storefronts describe the cohort. Included storefronts and assessed checkpoints describe usable policy-resolution evidence. Unknown and unsent checkpoints are not presumed successes or failures.

LaneCheckpoints: planned / attempted / observedRecorded submitted / assessed / unassessableAttained / unverifiedContexts: included / excludedStorefronts: included / registeredQuality eligible: contexts / storefrontsCaptures: original / repaired
Shopping250 / 183 / 183183 / 183 / 0104 / 3825 / 05 / 520 / 424 / 1
Support250 / 242 / 242242 / 242 / 096 / 13725 / 05 / 524 / 525 / 0
Study scope and limitations

Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.

This evaluation used the previously published method, frozen before capture and judging.

The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.

Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.

Unresolved capture failures pause publication; completed studies do not represent every deployment of a provider.

1 earlier submission attempt(s) preceded operator-corrected captures. They are retained in the repair provenance and excluded from the selected-capture scoring denominators.

Audit scope

Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible.

Provider names were masked where practicable; full anonymity is not claimed.

Each new quality judgment has a separate adversarial audit. The quality audit sees the primary verdict; the policy-resolution audit does not.