ALHENA RESEARCH LAB · VERSIONED SHOPPING FRAMEWORK

Shopping happens across the storefront.

shopping-journey-v2 evaluates what shoppers can accomplish and where they can access usable shopping assistance. It reports two separate scores: shopping outcomes and usable shopping reach. Neither is a replacement label for the existing chat composite.

Framework published; no results under this version yet.

The framework and evidence scoring rules are available. Automated capture across these interfaces is not yet enabled. Current submissions still run policy-resolution-v1. Existing scores have not been recalculated, and support scoring is unchanged.

Two dimensions, shown separately

More usable ways to access shopping assistance count toward shopping reach. Successful tasks count toward shopping outcomes. A tool can have broad reach and weak outcomes, or strong outcomes through fewer interfaces; the report keeps that distinction visible. There is no combined score or extra bonus for the number of widget instances.

Each evaluation registers five customer storefronts before testing. A study freezes its tasks, shopper constraints, primary routes, success conditions, observation flows and evidence requirements in advance and applies them consistently to each provider. Fallback attempts remain separate; the evaluator does not select a winning route after seeing results. The results describe those deployments, rather than every configuration the vendor offers.

Shopping outcomes: six equally weighted tasks

Six tasks per storefront, thirty task cells across five stores
TaskObservable outcome
Find a suitable productIdentify an available product satisfying every stated hard constraint, or correctly establish that none fits using current catalog evidence.
Answer product questions accuratelyGive the correct actionable answer grounded in the current product source; no fabricated or contradictory material claim.
Compare suitable alternativesExplain a verified meaningful difference and recommend a suitable option using the declared constraints.
Choose a purchasable variantSelect or direct the shopper to the exact available variant matching the request, verified on the product page.
Complete a cart actionActual cart state changes to the exact requested item, variant and quantity. A link, instruction or assistant claim without cart-state evidence fails. Never place an order.
Carry the shopping decision forwardThe chosen item, variant and relevant constraints survive the transition without forcing the shopper to restate them. Native product/cart destinations are allowed; additional provider widgets are not required.

Each assessable task earns binary credit for verified completion. A primary judgment and a separate audit must both support completion with valid evidence. A confirmed failure earns zero. A disagreement between the primary judgment and blind audit receives no verified completion credit and is disclosed. Unknown or blocked observations remain explicitly unresolved; they are not silently scored as zero or removed from the denominator.

A headline outcome score is available only when all thirty task cells are assessable and audited. It is 100 × verified completions ÷ 30, equivalent to averaging the six equally weighted tasks within each store and then averaging the five stores. Until then, the report shows coverage and the verified lower bound against all thirty planned tasks, not a completed headline score.

An assistant saying “added to cart” is insufficient. The cart-action task requires evidence of the actual cart state before and after the requested change. Continuity requires evidence that relevant context survives a transition; it does not require a vendor to offer additional widgets. Ordinary storefront navigation can be part of that transition. Testing stops before placing an order.

Usable shopping reach: five interface types

Each distinct interface type represents 20% of a storefront’s reach score
InterfaceWhat is inspected
Shopping chatA conversational entry point that provides usable shopping assistance.
Product-page Q&A / FAQProvider-powered product-specific assistance embedded on a product page. Static merchant FAQ copy alone does not count.
Embedded recommendationsProvider-powered interactive product recommendations outside the chat panel.
Search and discoveryProvider-powered shopping search or guided discovery outside the chat panel.
Cart assistanceA distinct provider-powered assistant embedded in or alongside the cart. A standard cart or a chat add-to-cart button alone is not another interface.

Discovery follows the homepage, a relevant category or search, a representative product page and a cart with a test item. Ordinary cookie consent and any blockers are recorded.

An interface counts only when it is attributable to the evaluated provider and a retained test demonstrates that a shopper can use it for its intended shopping purpose. Merely finding a button, widget or vendor marketing claim does not establish usability. Multiple placements of the same interface type count once per storefront.

Every interface observation retains its status and evidence
StatusMeaningScoring treatment
UsableAttribution and an actual usability test are verified.Earns credit for that type.
Present, untestedA surface was found but its required test is incomplete.Unresolved; no headline reach score.
Not observedThe interface was not observed in the sampled flow after the prescribed checks.No credit for that type in this sample.
BlockedAccess, attribution or observation could not be resolved.Unresolved; no headline reach score.

The reach score is the mean of the five storefront scores: 100 × verified usable interface cells ÷ 25. It is available only when every cell is resolved as usable or not observed. Otherwise, the score stays unavailable and the report shows the verified lower bound against all twenty-five cells alongside the unresolved inventory.

“Not observed” always means not observed in the sampled flow. It is not proof that a vendor lacks that capability. Product-page placement, the steps needed to open an interface and other interaction costs are recorded descriptively. This study does not establish actual adoption, increased conversion or a vendor’s intent or strategic focus.

Evidence, fresh runs and comparisons

Read the machine-readable framework and protocol hash. These task definitions and interface descriptions come from the same contract used by the evidence scorer.

Evidence identifies the storefront, page and interface; capture time; provider attribution; task and success condition; observed action and resulting state; and the primary judgment and audit. A report retains unknowns, failed tasks and capture limitations. Interface inventory remains distinct from vendor claims about available products.

Fresh shopping evidence is required for this version. Prior chat transcripts cannot establish product-page behavior, verified cart changes or continuity across interfaces. Compatible reuse remains limited to thirty days from original capture, with the same framework, task plan and execution requirements. Republishing does not reset evidence age.

Shopping journey scores are compared only with results from this same version under compatible conditions. They are not combined with prior shopping composites or used to recalculate support scores. Public summaries and the methodology remain open; detailed evidence requires a verified work email.

All research methodsPublished studies