Skip to content

Latest commit

 

History

History
82 lines (66 loc) · 4.05 KB

File metadata and controls

82 lines (66 loc) · 4.05 KB

Recommendation Evaluation Framework

Use this rubric to score AI shopping recommendations from 1 to 5. A recommendation should not be considered production-ready unless it scores at least 4 on relevance, price accuracy, availability, spec correctness, merchant reliability, and hallucination risk.

Score Definitions

Score Meaning
1 Fails the user need or presents unsafe/unverified claims
2 Partially useful, but important constraints or facts are missing
3 Reasonable draft recommendation with clear gaps
4 Good recommendation with verified facts and useful caveats
5 Excellent recommendation with strong evidence, fit, transparency, and conversion readiness

Dimensions

Dimension What To Check 1 Looks Like 5 Looks Like
Relevance Does the recommendation match the user's explicit need? Ignores category, use case, or budget Directly matches use case and constraints
Price Accuracy Are prices current, sourced, and bounded? Unsourced or stale prices Current offer, timestamp, merchant, and caveat
Availability Can the user actually buy it now? No stock signal In-stock status, delivery range, variant eligibility
Spec Correctness Are specs accurate and relevant? Incorrect or vague specs Verified specs from manufacturer or merchant feed
Review Trustworthiness Are reviews interpreted carefully? Treats star rating as truth Considers volume, recency, authenticity, and review distribution
Brand/Merchant Reliability Is the seller trustworthy? Unknown seller risk ignored Seller of record, return policy, warranty, fraud signals checked
Personal Fit Does it adapt to the user's preferences? Generic best-seller list Criteria-based ranking tied to user context
Explanation Quality Does the user understand the tradeoff? Shallow marketing copy Clear why, why not, and best-for guidance
Diversity Of Options Is the shortlist meaningfully varied? Near-duplicates Distinct budget, mainstream, upgrade, or use-case choices
Hallucination Risk Are unsupported claims avoided? Invented testing, prices, awards, or availability Facts separated from judgment, uncertainty exposed
Sponsored/Organic Clarity Are incentives clear? Paid ranking not disclosed Sponsored, affiliate, merchant, and organic signals labeled

Weighted Scoring

For most AI shopping journeys, use this weighting:

Dimension Weight
Relevance 15%
Price Accuracy 12%
Availability 12%
Spec Correctness 10%
Review Trustworthiness 8%
Brand/Merchant Reliability 10%
Personal Fit 10%
Explanation Quality 8%
Diversity Of Options 5%
Hallucination Risk 7%
Sponsored/Organic Clarity 3%

Production Gates

A recommendation should be blocked or downgraded if any of these are true:

  • Price is shown without merchant source or freshness.
  • Availability is unknown for a product framed as buyable.
  • The product does not satisfy a hard user constraint.
  • The answer claims hands-on testing without proof.
  • The recommendation relies on reviews but cannot explain review quality.
  • The seller is unknown or risky and no warning is shown.
  • A sponsored placement is blended into organic ranking.
  • The model invents specs, awards, or retailer policies.

Evaluation Template

Field Entry
Query
User constraints
Recommendation
Merchant
Price and timestamp
Availability
Key evidence
Main uncertainty
Score
Blocking issue?
What would improve it?

Example Read

For "best leakproof glass meal prep containers," a strong recommendation should not merely say "Pyrex is good." It should verify whether the specific set is glass, whether the lids are dishwasher safe, whether the containers are genuinely leakproof, whether the current offer is in stock, and whether a cheaper lookalike has review or seller risk.

The final answer should make the buyer's decision easier while making the system's uncertainty more visible.