Use this rubric to score AI shopping recommendations from 1 to 5. A recommendation should not be considered production-ready unless it scores at least 4 on relevance, price accuracy, availability, spec correctness, merchant reliability, and hallucination risk.
| Score | Meaning |
|---|---|
| 1 | Fails the user need or presents unsafe/unverified claims |
| 2 | Partially useful, but important constraints or facts are missing |
| 3 | Reasonable draft recommendation with clear gaps |
| 4 | Good recommendation with verified facts and useful caveats |
| 5 | Excellent recommendation with strong evidence, fit, transparency, and conversion readiness |
| Dimension | What To Check | 1 Looks Like | 5 Looks Like |
|---|---|---|---|
| Relevance | Does the recommendation match the user's explicit need? | Ignores category, use case, or budget | Directly matches use case and constraints |
| Price Accuracy | Are prices current, sourced, and bounded? | Unsourced or stale prices | Current offer, timestamp, merchant, and caveat |
| Availability | Can the user actually buy it now? | No stock signal | In-stock status, delivery range, variant eligibility |
| Spec Correctness | Are specs accurate and relevant? | Incorrect or vague specs | Verified specs from manufacturer or merchant feed |
| Review Trustworthiness | Are reviews interpreted carefully? | Treats star rating as truth | Considers volume, recency, authenticity, and review distribution |
| Brand/Merchant Reliability | Is the seller trustworthy? | Unknown seller risk ignored | Seller of record, return policy, warranty, fraud signals checked |
| Personal Fit | Does it adapt to the user's preferences? | Generic best-seller list | Criteria-based ranking tied to user context |
| Explanation Quality | Does the user understand the tradeoff? | Shallow marketing copy | Clear why, why not, and best-for guidance |
| Diversity Of Options | Is the shortlist meaningfully varied? | Near-duplicates | Distinct budget, mainstream, upgrade, or use-case choices |
| Hallucination Risk | Are unsupported claims avoided? | Invented testing, prices, awards, or availability | Facts separated from judgment, uncertainty exposed |
| Sponsored/Organic Clarity | Are incentives clear? | Paid ranking not disclosed | Sponsored, affiliate, merchant, and organic signals labeled |
For most AI shopping journeys, use this weighting:
| Dimension | Weight |
|---|---|
| Relevance | 15% |
| Price Accuracy | 12% |
| Availability | 12% |
| Spec Correctness | 10% |
| Review Trustworthiness | 8% |
| Brand/Merchant Reliability | 10% |
| Personal Fit | 10% |
| Explanation Quality | 8% |
| Diversity Of Options | 5% |
| Hallucination Risk | 7% |
| Sponsored/Organic Clarity | 3% |
A recommendation should be blocked or downgraded if any of these are true:
- Price is shown without merchant source or freshness.
- Availability is unknown for a product framed as buyable.
- The product does not satisfy a hard user constraint.
- The answer claims hands-on testing without proof.
- The recommendation relies on reviews but cannot explain review quality.
- The seller is unknown or risky and no warning is shown.
- A sponsored placement is blended into organic ranking.
- The model invents specs, awards, or retailer policies.
| Field | Entry |
|---|---|
| Query | |
| User constraints | |
| Recommendation | |
| Merchant | |
| Price and timestamp | |
| Availability | |
| Key evidence | |
| Main uncertainty | |
| Score | |
| Blocking issue? | |
| What would improve it? |
For "best leakproof glass meal prep containers," a strong recommendation should not merely say "Pyrex is good." It should verify whether the specific set is glass, whether the lids are dishwasher safe, whether the containers are genuinely leakproof, whether the current offer is in stock, and whether a cheaper lookalike has review or seller risk.
The final answer should make the buyer's decision easier while making the system's uncertainty more visible.