This document describes the agent skills available in the Evret project. These skills are used by Claude Code to provide specialized guidance on retrieval evaluation metrics and RAG pipeline development.
.agents/
└── skills/
└── retriever-evals/
├── SKILL.md # Main RAG metrics skill
└── metrics/
├── HITRATE_SKILL.md # Hit Rate metric guidance
├── RECALL_SKILL.md # Recall@K metric guidance
├── PRECISION_SKILL.md # Precision@K metric guidance
├── MRR_SKILL.md # MRR metric guidance
├── NDCG_SKILL.md # nDCG metric guidance
├── ERR_SKILL.md # ERR metric guidance
├── RBP_SKILL.md # RBP metric guidance
└── AVERAGE_PRECISION_SKILL.md # Average Precision/MAP guidance
Location: .agents/skills/retriever-evals/SKILL.md
Trigger Keywords:
- RAG pipeline, retrieval system, vector search
- MRR, nDCG, Hit Rate, Recall, retrieval metrics, ranking metrics
- Debugging hallucinating LLM
- Tuning retriever or reranker
- Choosing between retrieval evaluation metrics
Purpose: Provides high-level guidance on choosing the right retrieval evaluation metric for different use cases.
Key Topics:
- Metric selection guidance based on use case
- Stage-based recommendations (debugging, tuning, eval)
- Comparison table of all metrics
- Best practices for multi-objective evaluation
Quick Reference:
| Use Case | Recommended Metric | Why |
|---|---|---|
| Check if any correct result retrieved | Hit Rate | Binary success check |
| Measure completeness of results | Recall | Coverage metric |
| Know how early first correct result appears | MRR | Single-answer QA optimization |
| Care about ranking quality | nDCG | Multi-relevance ranking quality |
| Model user satisfaction with graded relevance | ERR | Cascade browsing behavior |
| Model result quality with user patience | RBP | Tunable position weighting |
Location: .agents/skills/retriever-evals/metrics/HITRATE_SKILL.md
Trigger Keywords:
- Hit Rate
- Binary check on retrieval
- Debugging hallucinating LLM
- Simplest retrieval metric
- Validating chunks in top-k results
Purpose: Deep dive into Hit Rate metric - the first-line diagnostic for RAG systems.
Formula:
Hit Rate = (1/|Q|) × Σ 𝟙[relevant ∩ retrieved ≠ ∅]
When to Use:
- ✅ Initial RAG debugging / baseline check
- ✅ Embedding model or chunking strategy swap
- ✅ Reranker evaluation (alongside nDCG)
- ❌ Evaluating ranking quality (use MRR or nDCG)
- ❌ Completeness-critical retrieval (use Recall)
Best Practices:
- Use as entry-point metric, not final one
- Choose k that mirrors production context window
- Segment by query difficulty and domain
- Pair with MRR to separate presence from position
- Use in CI/CD to catch regressions
Location: .agents/skills/retriever-evals/metrics/RECALL_SKILL.md
Trigger Keywords:
- Recall, Recall@K
- Coverage metric
- Completeness of retrieval
- Multi-document QA
Purpose: Measures what fraction of all relevant documents were retrieved.
Formula:
Recall@k = |relevant ∩ retrieved_k| / |relevant|
When to Use:
- ✅ Multi-document QA or synthesis tasks
- ✅ Completeness-critical domains (medical, financial, safety)
- ✅ Measuring coverage improvements
- ❌ When you only care about first hit (use MRR)
- ❌ When ranking quality matters more (use nDCG)
Location: .agents/skills/retriever-evals/metrics/PRECISION_SKILL.md
Trigger Keywords:
- Precision, Precision@K
- Purity metric
- LLM context quality
- Noise reduction
Purpose: Measures what fraction of retrieved documents are relevant.
Formula:
Precision@k = |relevant ∩ retrieved_k| / k
When to Use:
- ✅ LLM context quality optimization
- ✅ Noise reduction in retrieved results
- ✅ Limited context window scenarios
- ❌ When you need to measure coverage (use Recall)
- ❌ When ranking order matters (use MRR or nDCG)
Location: .agents/skills/retriever-evals/metrics/MRR_SKILL.md
Trigger Keywords:
- MRR, Mean Reciprocal Rank
- First-hit metric
- Single-answer QA
- Fast-hit evaluation
Purpose: Measures how early the first relevant result appears in rankings.
Formula:
MRR = (1/|Q|) × Σ (1/rank_first_relevant)
Scoring:
- Rank 1 → score = 1.0
- Rank 2 → score = 0.5
- Rank 5 → score = 0.2
When to Use:
- ✅ Single-answer QA systems
- ✅ Intent classification, slot filling
- ✅ Fast-hit optimization
- ❌ Multi-document tasks (use nDCG)
- ❌ When you need all relevant results (use Recall)
Location: .agents/skills/retriever-evals/metrics/NDCG_SKILL.md
Trigger Keywords:
- nDCG, DCG, Normalized Discounted Cumulative Gain
- Ranking quality
- Graded relevance
- Reranker tuning
Purpose: Measures ranking quality across all relevant results with support for graded relevance.
Formula:
DCG@k = Σ (rel_i / log₂(i+1))
nDCG@k = DCG@k / IDCG@k
Positional Discount:
- Rank 1: discount = 1.0
- Rank 2: discount ≈ 0.63
- Rank 5: discount ≈ 0.39
When to Use:
- ✅ Reranker evaluation
- ✅ Multi-document QA or synthesis
- ✅ Search systems with graded relevance
- ✅ System-level RAG evaluation
- ❌ Simple RAG debugging (use Hit Rate)
- ❌ Single-answer QA (use MRR)
Best Practices:
- Invest in graded relevance labels (exact/partial/irrelevant)
- Match k to reranker's output size
- Pair with Hit Rate to prevent coverage regression
- Compare against BM25 baseline
- Normalize relevance scale across annotators
Location: .agents/skills/retriever-evals/metrics/ERR_SKILL.md
Trigger Keywords:
- ERR, Expected Reciprocal Rank
- Cascade model
- User satisfaction
- Graded relevance
Purpose: Models how likely a user is to be satisfied while scanning ranked results from top to bottom.
Formula:
ERR@k = Σ(i=1 to k) [(1/i) × R(i) × Π(j=1 to i-1)(1 - R(j))]
R(i) = (2^grade(i) - 1) / 2^max_grade
When to Use:
- ✅ Graded relevance judgments are available
- ✅ User satisfaction and stopping behavior matter
- ✅ Earlier highly relevant results should dominate the score
- ❌ Binary-only first-hit checks (use MRR or Hit Rate)
- ❌ Coverage-critical retrieval (use Recall)
Location: .agents/skills/retriever-evals/metrics/RBP_SKILL.md
Trigger Keywords:
- RBP, Rank-Biased Precision
- Persistence parameter
- User patience
- Position-weighted precision
Purpose: Measures ranked result quality using a tunable persistence parameter that controls how deep users are expected to look.
Formula:
RBP(p)@k = (1 - p) × Σ(i=1 to k) [p^(i-1) × rel(i)]
When to Use:
- ✅ User patience differs by product or search mode
- ✅ You need tunable position weighting
- ✅ Comparing incomplete or different-length rankings
- ❌ You need all relevant documents recovered (use Recall)
- ❌ You need cascade satisfaction behavior (use ERR)
Location: .agents/skills/retriever-evals/metrics/AVERAGE_PRECISION_SKILL.md
Trigger Keywords:
- Average Precision, AP, MAP
- Mean Average Precision
- Rank-aware precision
- Benchmark comparison
Purpose: Rank-aware precision metric that considers all relevant results and their positions.
Formula:
AP = (1/R) × Σ P(k) × 𝟙[doc_k is relevant]
Where R is the total number of relevant documents.
When to Use:
- ✅ Benchmark comparison across systems
- ✅ Multi-relevant query evaluation
- ✅ Academic/research contexts
- ❌ When only first hit matters (use MRR)
- ❌ When graded relevance is needed (use nDCG)
Start: What are you trying to optimize?
├─ Basic functionality check?
│ └─ Use: Hit Rate
│
├─ How early is the first relevant result?
│ └─ Use: MRR
│
├─ How many relevant results did we find?
│ └─ Use: Recall
│
├─ How pure are the results?
│ └─ Use: Precision
│
├─ How well are results ranked? (binary relevance)
│ └─ Use: Average Precision (MAP)
│
├─ How well are results ranked? (graded relevance)
│ └─ Use: nDCG
│
├─ How satisfied is a user while scanning results?
│ └─ Use: ERR
│
└─ How does ranking quality change with user patience?
└─ Use: RBP
These skills guide the implementation and usage of the metrics in the Evret framework:
Implementation Reference:
- Metric base class: src/evret/metrics/base.py
- Individual metrics: src/evret/metrics/
- Evaluation orchestrator: src/evret/evaluation/evaluator.py
Usage Example:
from evret.retrievers import QdrantRetriever
from evret.evaluation import Evaluator, Dataset
from evret.metrics import ERR, HitRate, MRR, NDCG, RBP
# Setup retriever
retriever = QdrantRetriever(
url="http://localhost:6333",
collection_name="my_docs"
)
# Load evaluation dataset
dataset = Dataset.from_json("eval_data.json")
# Run evaluation with multiple metrics
evaluator = Evaluator(
retriever=retriever,
metrics=[
HitRate(k=5), # Binary presence check
MRR(k=10), # First-hit optimization
NDCG(k=5), # Ranking quality
ERR(k=5), # User satisfaction
RBP(k=5), # User patience weighting
]
)
results = evaluator.evaluate(dataset)
print(results.summary())| Metric | Type | Rank-Sensitive | First Hit Only | Best Use Case |
|---|---|---|---|---|
| Hit Rate | Binary | No (threshold) | No | RAG debugging, presence check |
| Recall | Binary/Graded | No | No | Completeness, coverage |
| Precision | Binary/Graded | No | No | Context quality, noise reduction |
| MRR | Binary | Yes | Yes | Fast hits, single-answer QA |
| nDCG | Graded | Yes | No | Ranking quality, search systems |
| ERR | Graded | Yes | No | User satisfaction, cascade behavior |
| RBP | Binary/Graded | Yes | No | User patience, position-weighted quality |
| Average Precision | Binary | Yes | No | Benchmark comparison, multi-relevant |
To add a new skill to the Evret agent system:
-
Create a new
.mdfile in .agents/skills/retriever-evals/metrics/ (for metric skills) or .agents/skills/retriever-evals/ (for general skills) -
Follow the frontmatter format:
--- name: skill-name description: Trigger conditions and use cases ---
-
Include these sections:
- What is the use of this metric/skill
- Mathematical Formula (for metrics)
- When to use this metric/skill
- Best practices
-
Update this
Agents.mdfile with the new skill documentation
- Project prompt: PROMPT.md
- Project progress: progress.md
- Contributing guide: CONTRIBUTING.md
- Main README: README.md