THE METHOD

A fixed rubric.
An open evidence record.

Alhena Research Lab applies Gorgias’s published shopping and support quality criteria. The weights stay fixed, and every decision is backed by a captured response.

One tool at a time

One tool, three selected customer storefronts. One ten-turn shopping conversation and one ten-turn returns conversation per storefront: six conversations, 60 turns and 78 criterion decisions per tool.

Separate judging and audit

An AI judge scores each criterion. A separate AI audit reviews those decisions. Code calculates the points using the published weights and signal requirements.

Fresh browser sessions

Newly captured conversations use a fresh browser context. An unsupported widget, a human takeover or incomplete evidence stops the run. The original pilot’s session limitations remain disclosed.

Validation before publication

Approved runs publish automatically only after checking completeness, evidence references, score calculations and audit coverage. Failed runs remain private; missing answers do not become invented scores.

This is a quality pilot, not the full Gorgias benchmark or its composite leaderboard. It uses the everyday-value and returns themes. No automation or speed ranking is claimed. Three conversations per provider per lane do not meet the original benchmark’s 15-conversation threshold. Selected storefronts and deployment configurations limit generalization.

Shopping quality

16 criteria · 100 available points

  • Answer quality30 pts
  • Discovery20 pts
  • Recommendation22 pts
  • Rich elements18 pts
  • Closing10 pts

Support quality

10 criteria · 100 available points

  • Resolution40 pts
  • Accuracy25 pts
  • Actionability20 pts
  • Closing15 pts

Shopping · 16 criteria · 100 points

CriterionPointsPublished pass condition
a_direct
answer
14each shopper question gets a direct, on-topic, **substantive** response. A generic greeting or brand boilerplate ("Hi! How can I help?", "We're so happy to have you here!") that does not actually answer the question **fails** — warmth, enthusiasm and emojis are not substance.
a_consistent
answer
9no contradiction across turns; no invented policies/specs
a_no_ignored
answer
7no shopper turn is left unanswered or answered with an unrelated reply
d_clarify
discovery
8on a BROAD/ambiguous opener ("a gift", "where do I start", "something for X") it asks a **relevant** clarifying question (budget, recipient, use-case, skin type, size, preference) before recommending. On an already-specific ask, going straight to a targeted answer also passes — the fail is *guessing blind on a vague ask*.
d_progressive
discovery
7it **uses the shopper's answers** to narrow/refine across turns (builds a profile) — a later recommendation reflects earlier stated constraints, not a generic pick repeated.
d_not_dump
discovery
5it does **not** dump a generic best-seller list on a vague opener before understanding needs (a wall of unqualified products fails).
r_named
recommendation
9recommends at least one specific, named product
r_fit
recommendation
8ties the recommendation to the shopper's **specific stated constraints** (budget, use-case, skin type, recipient…) with real reasoning. A product dropped with generic praise/emojis but no rationale connected to what the shopper actually said **fails**.
r_plausible
recommendation
5recommendations are appropriate to the store's catalog and the request
e_price
rich
6a concrete price is attached to a recommended product

Required signal: has_price

e_link
rich
7a product link/card the shopper can open is presented

Required signal: has_link

e_reviews
rich
3review counts/ratings (or images) support the recommendation

Required signal: has_reviews

e_options
rich
2multiple distinct options are laid out for comparison

Required signal: has_options

c_cta
close
5ends with a concrete next step (open the product, pick a size, apply code…)
c_cart
close
3facilitates purchase mechanics (add-to-cart, checkout guidance). If the shopper **explicitly** asks to add to cart or for a total/checkout and the agent does not actually do it (re-asks, deflects, or loops), this **fails** — an unfulfilled explicit purchase request is a hard miss.
c_clean
close
2conversation closes cleanly (no dangling question, no mid-thought stop)

Support · 10 criteria · 100 points

CriterionPointsPublished pass condition
s_answered
resolution
18the PRIMARY ask gets a complete, store-specific answer or procedure in-channel. A generic pointer ("check our returns page"), a partial answer, or industry-generality boilerplate fails

Required signal: no_deflect

s_outcome
resolution
12the ending is outcome-correct: resolved in-channel, **or** a justified handover done well (context collected first, expectations set). An unjustified bail on an answerable question fails

Required signal: no_deflect

s_no_deflect
resolution
10doesn't push the shopper out of channel ("email us", "call us") when the question was answerable in-chat

Required signal: no_deflect

g_specific
accuracy
13gives ≥2 concrete, store-specific facts (a number, timeframe, named policy term, concrete condition). One vague reassurance ("we'll sort it out") fails
g_consistent
accuracy
5no self-contradiction across turns
g_grounded
accuracy
7policy/product claims read as grounded in THIS store (named policies, actual conditions, store-specific procedures) rather than plausible industry generalities
t_steps
actionability
12the shopper leaves with steps they can execute NOW in their situation (cold session, no account) — numbered or clearly sequenced; "reach out if…" alone fails
t_complete
actionability
8the answer covers the actual ask (not a fragment of it)
k_expectations
close
8sets expectations with WHO acts and WHEN (a timeframe) — vague "soon"/"we'll be in touch" fails
k_clean
close
7clean, complete close

Source and version

The source rubric is published by Gorgias. Alhena Research Lab is an Alhena-operated project and is not endorsed or certified by Gorgias. Read the canonical rubric.

Pinned source commit: 19b1420d2520d48baa52be81ac33fc4b9bd0ff8b. Protocol: quality-pilot-v1.

What a score does not prove

The tool library shows one complete three-storefront evaluation for each tool. After a new evaluation passes validation, comparison reports are assembled against other compatible, recent tool evaluations. Each pair contains 12 source conversations and 120 captured turns; composing the report does not run new conversations or judge them again. Repeated use in comparison reports does not increase a tool’s sample size.

Compatible analysis may be reused for 30 days from its original capture date. Reused conversations link to their source report and retain their original dates and limitations. Republishing does not reset this window. Only missing or expired conversations are evaluated again.

The imported September 20 study did not record per-turn AI-author proof. Its attribution is inherited from the original report when reused; no new author verification is performed. New automated captures require positive AI-author evidence before publication.

A high score describes how these captured conversations satisfied the specified rubric. It does not establish that a provider is universally better, that all factual claims are correct, or that a storefront will produce the same responses every time.

Separate AI judging and auditing are parts of this commissioned evaluation. They do not make Alhena Research Lab an independent research institution. Full transcripts, scoring details, source checks and limitations accompany each completed report. Summaries and this rubric are public; detailed evidence requires a verified work email.