THE METHOD
A fixed rubric.
An open evidence record.
Alhena Research Lab applies Gorgias’s published shopping and support quality criteria. The weights stay fixed, and every decision is backed by a captured response.
One tool at a time
One tool, three selected customer storefronts. One ten-turn shopping conversation and one ten-turn returns conversation per storefront: six conversations, 60 turns and 78 criterion decisions per tool.
Separate judging and audit
An AI judge scores each criterion. A separate AI audit reviews those decisions. Code calculates the points using the published weights and signal requirements.
Fresh browser sessions
Newly captured conversations use a fresh browser context. An unsupported widget, a human takeover or incomplete evidence stops the run. The original pilot’s session limitations remain disclosed.
Validation before publication
Approved runs publish automatically only after checking completeness, evidence references, score calculations and audit coverage. Failed runs remain private; missing answers do not become invented scores.
This is a quality pilot, not the full Gorgias benchmark or its composite leaderboard. It uses the everyday-value and returns themes. No automation or speed ranking is claimed. Three conversations per provider per lane do not meet the original benchmark’s 15-conversation threshold. Selected storefronts and deployment configurations limit generalization.
Shopping quality
16 criteria · 100 available points
- Answer quality30 pts
- Discovery20 pts
- Recommendation22 pts
- Rich elements18 pts
- Closing10 pts
Support quality
10 criteria · 100 available points
- Resolution40 pts
- Accuracy25 pts
- Actionability20 pts
- Closing15 pts
Shopping · 16 criteria · 100 points
| Criterion | Points | Published pass condition |
|---|---|---|
| a_direct answer | 14 | each shopper question gets a direct, on-topic, **substantive** response. A generic greeting or brand boilerplate ("Hi! How can I help?", "We're so happy to have you here!") that does not actually answer the question **fails** — warmth, enthusiasm and emojis are not substance. |
| a_consistent answer | 9 | no contradiction across turns; no invented policies/specs |
| a_no_ignored answer | 7 | no shopper turn is left unanswered or answered with an unrelated reply |
| d_clarify discovery | 8 | on a BROAD/ambiguous opener ("a gift", "where do I start", "something for X") it asks a **relevant** clarifying question (budget, recipient, use-case, skin type, size, preference) before recommending. On an already-specific ask, going straight to a targeted answer also passes — the fail is *guessing blind on a vague ask*. |
| d_progressive discovery | 7 | it **uses the shopper's answers** to narrow/refine across turns (builds a profile) — a later recommendation reflects earlier stated constraints, not a generic pick repeated. |
| d_not_dump discovery | 5 | it does **not** dump a generic best-seller list on a vague opener before understanding needs (a wall of unqualified products fails). |
| r_named recommendation | 9 | recommends at least one specific, named product |
| r_fit recommendation | 8 | ties the recommendation to the shopper's **specific stated constraints** (budget, use-case, skin type, recipient…) with real reasoning. A product dropped with generic praise/emojis but no rationale connected to what the shopper actually said **fails**. |
| r_plausible recommendation | 5 | recommendations are appropriate to the store's catalog and the request |
| e_price rich | 6 | a concrete price is attached to a recommended product Required signal: has_price |
| e_link rich | 7 | a product link/card the shopper can open is presented Required signal: has_link |
| e_reviews rich | 3 | review counts/ratings (or images) support the recommendation Required signal: has_reviews |
| e_options rich | 2 | multiple distinct options are laid out for comparison Required signal: has_options |
| c_cta close | 5 | ends with a concrete next step (open the product, pick a size, apply code…) |
| c_cart close | 3 | facilitates purchase mechanics (add-to-cart, checkout guidance). If the shopper **explicitly** asks to add to cart or for a total/checkout and the agent does not actually do it (re-asks, deflects, or loops), this **fails** — an unfulfilled explicit purchase request is a hard miss. |
| c_clean close | 2 | conversation closes cleanly (no dangling question, no mid-thought stop) |
Support · 10 criteria · 100 points
| Criterion | Points | Published pass condition |
|---|---|---|
| s_answered resolution | 18 | the PRIMARY ask gets a complete, store-specific answer or procedure in-channel. A generic pointer ("check our returns page"), a partial answer, or industry-generality boilerplate fails Required signal: no_deflect |
| s_outcome resolution | 12 | the ending is outcome-correct: resolved in-channel, **or** a justified handover done well (context collected first, expectations set). An unjustified bail on an answerable question fails Required signal: no_deflect |
| s_no_deflect resolution | 10 | doesn't push the shopper out of channel ("email us", "call us") when the question was answerable in-chat Required signal: no_deflect |
| g_specific accuracy | 13 | gives ≥2 concrete, store-specific facts (a number, timeframe, named policy term, concrete condition). One vague reassurance ("we'll sort it out") fails |
| g_consistent accuracy | 5 | no self-contradiction across turns |
| g_grounded accuracy | 7 | policy/product claims read as grounded in THIS store (named policies, actual conditions, store-specific procedures) rather than plausible industry generalities |
| t_steps actionability | 12 | the shopper leaves with steps they can execute NOW in their situation (cold session, no account) — numbered or clearly sequenced; "reach out if…" alone fails |
| t_complete actionability | 8 | the answer covers the actual ask (not a fragment of it) |
| k_expectations close | 8 | sets expectations with WHO acts and WHEN (a timeframe) — vague "soon"/"we'll be in touch" fails |
| k_clean close | 7 | clean, complete close |
Source and version
The source rubric is published by Gorgias. Alhena Research Lab is an Alhena-operated project and is not endorsed or certified by Gorgias. Read the canonical rubric.
Pinned source commit: 19b1420d2520d48baa52be81ac33fc4b9bd0ff8b. Protocol: quality-pilot-v1.
What a score does not prove
The tool library shows one complete three-storefront evaluation for each tool. After a new evaluation passes validation, comparison reports are assembled against other compatible, recent tool evaluations. Each pair contains 12 source conversations and 120 captured turns; composing the report does not run new conversations or judge them again. Repeated use in comparison reports does not increase a tool’s sample size.
Compatible analysis may be reused for 30 days from its original capture date. Reused conversations link to their source report and retain their original dates and limitations. Republishing does not reset this window. Only missing or expired conversations are evaluated again.
The imported September 20 study did not record per-turn AI-author proof. Its attribution is inherited from the original report when reused; no new author verification is performed. New automated captures require positive AI-author evidence before publication.
A high score describes how these captured conversations satisfied the specified rubric. It does not establish that a provider is universally better, that all factual claims are correct, or that a storefront will produce the same responses every time.
Separate AI judging and auditing are parts of this commissioned evaluation. They do not make Alhena Research Lab an independent research institution. Full transcripts, scoring details, source checks and limitations accompany each completed report. Summaries and this rubric are public; detailed evidence requires a verified work email.