A study of shopping and support on ten storefront deployments, with conversation evidence, public merchant policies, fixed quality criteria and separate response timing.
Commissioned by Alhena Research Lab. Published Sep 22, 2026. Captures Sep 21, 2026 to Sep 22, 2026.
Interpret the scores within their public-session scope.
Policy-compliant resolution can include a verified policy-required next step. It does not establish actual refunds, completed account actions or human-resolved orders. This is a separately versioned Alhena Research Lab method, not an independent certification or the Gorgias leaderboard.
Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.
This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments.
The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.
Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.
All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates.
Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained.
Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established.
The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed.
Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores.
Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
Quality
83 / 100
The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
Full-answer speed score
53.7 / 100
Full-answer completion 11.8 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
Quality
72 / 100
The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
Full-answer speed score
25.3 / 100
Full-answer completion 17.2 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
Quality
95 / 100
The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
Full-answer speed score
70 / 100
Full-answer completion 8.7 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.
Quality
94 / 100
The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.
Full-answer speed score
41.6 / 100
Full-answer completion 14.1 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.
support evidence coverage
Company
Checkpoints: planned / attempted / observed
Recorded submitted / assessed / unassessable
Attained / unverified
Contexts: included / excluded
Stores: included / excluded
Quality eligible: contexts / stores
Captures: original / repaired
Alhena
250 / 250 / 250
250 / 250 / 0
182 / 63
25 / 0
5 / 0
25 / 5
19 / 6
Gorgias
250 / 205 / 205
205 / 180 / 25
122 / 45
18 / 7
4 / 1
18 / 4
23 / 2
Overall composite
Alhena
Overall composite
74.9 / 100
Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.
Gorgias
Overall composite
65.8 / 100
Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.
Method and audit
Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible.
Provider names were masked where practicable; full anonymity is not claimed.
Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit.
One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts.
The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance.
A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts.
The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance.
The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage.
Replaces the source automation classifier with policy-compliant resolution.
Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence.
Cause-selected capture repairs retained: 12. The source study remains unchanged.
Detailed decisions, source records and available study downloads require a verified work email. Your email and view or download activity are shared privately with Alhena. Privacy details.