All policy-resolution studies

ALHENA RESEARCH LAB · policy-resolution-v1

Alhena and Gorgias: policy-compliant resolution

A study of shopping and support on ten storefront deployments, with conversation evidence, public merchant policies, fixed quality criteria and separate response timing.

Commissioned by Alhena Research Lab. Published Sep 22, 2026. Captures Sep 21, 2026 to Sep 22, 2026.

Interpret the scores within their public-session scope.

Policy-compliant resolution can include a verified policy-required next step. It does not establish actual refunds, completed account actions or human-resolved orders. This is a separately versioned Alhena Research Lab method, not an independent certification or the Gorgias leaderboard.

Alhena commissioned and operates this study. Selected storefront deployments do not establish a universal provider ranking or independent certification.

This methodology was developed after reviewing the original automation results, then frozen before the new policy-resolution judgments.

The sessions were logged out. No real order, refund, account change or completed human resolution was independently verified.

Unknown delivery or actor attribution remains unassessable. Unsent questions are not successes or failures. Conditional scores must be read with planned and assessed coverage.

All ten JSHealth core contexts were excluded because a generative actor and delivery could not be confirmed. Two additional support contexts stopped at login gates.

Two retained dryrobe shopping contexts have only six measured replies. Twelve cause-selected capture repairs contain all 120 submitted questions; every repaired outcome is retained.

Merchant policy conflicts and gaps are retained. Product efficacy, review authenticity and downstream cart or refund completion were not independently established.

The repaired captures have gaps in live resource telemetry. Terminal state and artifact hashes were verified; continuous resource observations are not claimed.

Private cart, session or customer links are redacted from displayed evidence and quotations. Judgments used the unchanged original captures; these display redactions do not change scores.

100Captured core contexts
88Judged core contexts
10Guardrail contexts
836Audited PCR decisions

Shopping

Alhena

Composite

68.4 / 100

40% policy-compliant resolution, 35% quality, 25% speed.

Policy-compliant resolution · public-session scope

64.8 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Quality

83 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

53.7 / 100

Full-answer completion 11.8 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

Gorgias

Composite

55.8 / 100

40% policy-compliant resolution, 35% quality, 25% speed.

Policy-compliant resolution · public-session scope

60.8 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Quality

72 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

25.3 / 100

Full-answer completion 17.2 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

shopping evidence coverage
CompanyCheckpoints: planned / attempted / observedRecorded submitted / assessed / unassessableAttained / unverifiedContexts: included / excludedStores: included / excludedQuality eligible: contexts / storesCaptures: original / repaired
Alhena250 / 242 / 242242 / 242 / 0158 / 6425 / 05 / 025 / 523 / 2
Gorgias250 / 187 / 187187 / 164 / 23109 / 3720 / 54 / 117 / 423 / 2

Support

Alhena

Composite

81.4 / 100

50% policy-compliant resolution, 40% quality, 10% speed.

Policy-compliant resolution · public-session scope

72.8 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Quality

95 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

70 / 100

Full-answer completion 8.7 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

Gorgias

Composite

75.8 / 100

50% policy-compliant resolution, 40% quality, 10% speed.

Policy-compliant resolution · public-session scope

68 / 100

Verified correct answer or policy-prescribed next step, averaged equally by conversation and storefront. This does not measure downstream case completion.

Quality

94 / 100

The original fixed answer-quality rubric. Unchanged captures retain their original judgments; repaired captures receive a fresh judgment and audit.

Full-answer speed score

41.6 / 100

Full-answer completion 14.1 seconds; 100 at 3 seconds, zero at 22 seconds, bounded between 0 and 100.

support evidence coverage
CompanyCheckpoints: planned / attempted / observedRecorded submitted / assessed / unassessableAttained / unverifiedContexts: included / excludedStores: included / excludedQuality eligible: contexts / storesCaptures: original / repaired
Alhena250 / 250 / 250250 / 250 / 0182 / 6325 / 05 / 025 / 519 / 6
Gorgias250 / 205 / 205205 / 180 / 25122 / 4518 / 74 / 118 / 423 / 2

Overall composite

Alhena

Overall composite

74.9 / 100

Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.

Gorgias

Overall composite

65.8 / 100

Equal mean of the two unrounded lane composites; only available when both lanes meet coverage floors.

Method and audit

Every assessed policy-resolution checkpoint receives a primary and a fresh blind audit. Both must award valid evidence-backed attainment. Attainment disagreements receive zero verified credit and remain visible.

Provider names were masked where practicable; full anonymity is not claimed.

Retained original quality judgments have separate historical audit coverage. The repair-quality audit sees the primary quality verdicts and full captured replies; it is distinct from the blind policy-resolution audit.

One initial scoring attempt failed in the response decoder with zero accepted decisions; upstream completion is unknown. The output was unavailable, and the documented infrastructure repair did not select among verdicts.

The first blind policy-resolution audit exhausted its 8,000-token per-model-turn output allowance without a structured verdict. Its unusable output was preserved and the identical input was audited once more under the reviewed 16,384-token per-model-turn allowance.

A later blind audit failed after about 602 seconds without a returned verdict. This is consistent with the original 600-second local timeout, but the inner error was not retained and provider completion is unknown. The unusable attempt was preserved and the identical input was retried once under a uniform 1,200-second future execution limit with a 1,250-second outer transport limit. All 36 previously successful stage outputs remain unchanged under their original timeout contracts.

The exact successful initial primary judgment for 30 checkpoints is retained with an 8,000-token per-model-turn allowance. Every subsequent policy-resolution and repair-quality call uses 16,384 per model turn, with the same frozen prompts, model, effort and criteria. New calls use fresh independent sessions under reviewed pool releases with 2 and 4 active slots; 4 distinct isolated slots were used. A CLI session can contain multiple internal steps, so its total output can exceed the per-turn allowance.

The original quality audit sampled only Alhena records. New policy-resolution audits cover both providers, but this does not retroactively expand historical quality-audit coverage.

Replaces the source automation classifier with policy-compliant resolution.

Every eligible checkpoint receives two fresh blind judgments; attainment requires agreement and valid evidence.

Cause-selected capture repairs retained: 12. The source study remains unchanged.

Read the versioned methodology

Frozen method SHA-256: 712b1e5753cd9b871881d328d79ecb20ccbb279c045fe9233891c6805441133b

The evidence behind the scores

Detailed decisions, source records and available study downloads require a verified work email. Your email and view or download activity are shared privately with Alhena. Privacy details.