mgq-retrieval-report evaluates any TREC-format run.tsv against a qrels
file and writes reproducible JSON and Markdown artifacts. It is intended for
BM25, dense, hybrid RRF, and reranked runs, as long as each input follows the
standard six-column run format:
qid Q0 doc_id rank score system
Use evaluate when a run should be scored independently:
mgq-retrieval-report evaluate \
--run outputs/dense_retrieval/run.tsv \
--run-name dense \
--output-dir outputs/retrieval_reports/denseArtifacts:
metrics.json: input run/qrels paths, metric values, metric settings, evaluated qid count, and skipped-qid coverage diagnostics.report.md: a compact Markdown summary for experiment notes.
By default the report computes MRR@10, nDCG@10, Recall@100, and
Recall@1000. Override cutoffs with --ks-mrr, --ks-ndcg, and
--ks-recall.
Use compare when the claim is comparative. The command restricts both runs
to the same qid set with positive qrels before computing deltas, which avoids
misleading comparisons when two retrieval stages cover different queries.
mgq-retrieval-report compare \
--baseline-run outputs/dense_retrieval/run.tsv \
--candidate-run outputs/hybrid_rrf/run.tsv \
--baseline-name dense \
--candidate-name rrf \
--output-dir outputs/retrieval_reports/dense_vs_rrfArtifacts:
comparison.json: input run/qrels paths, matched metrics, candidate-minus-baseline deltas, coverage counts, and movement-bucket summary.per_query.jsonl: query-level promoted, demoted, new-hit, lost-hit, unchanged-hit, and unchanged-miss diagnostics.report.md: a compact Markdown summary.
The coverage block is part of the result, not a footnote. Check
n_baseline_only_qids, n_candidate_only_qids, and
n_shared_without_positive_qrels before interpreting metric deltas.
Use matrix when a research claim compares more than two systems, such as
BM25-on-sample, dense, RRF, and RRF followed by cross-encoder reranking. The
command restricts every row to qids shared by all input runs and positive
qrels before computing metrics:
mgq-retrieval-report matrix \
--run bm25_sample=outputs/dense_retrieval/run_bm25_sample.tsv \
--run dense=outputs/dense_retrieval/run.tsv \
--run rrf=outputs/hybrid_rrf/run.tsv \
--run rrf_reranked=outputs/hybrid_rrf_rerank/run.tsv \
--baseline-name bm25_sample \
--output-dir outputs/retrieval_reports/hybrid_matrixArtifacts:
matrix.json: run order, per-run metrics, candidate-minus-baseline deltas, best run per metric, shared-qid coverage, and movement diagnostics.pairwise_deltas.jsonl: one row per candidate vs the named baseline.report.md: Markdown tables for the metrics matrix, deltas, best run per metric, and coverage.
Prefer this command over hand-copying metrics from separate runs. Separate
metrics.json files may cover different query sets; the matrix report makes
the shared-qid restriction explicit and repeatable.
If --qrels is omitted, the command loads MS MARCO Passage dev/small qrels
through ir_datasets. Pass --qrels when evaluating a custom split,
sample-specific qrels file, or downloaded TREC qrels artifact.
The command accepts standard four-column TREC qrels:
qid iter doc_id relevance
It also accepts compact three-column qrels:
qid doc_id relevance
Only positive relevance labels count as relevant. Non-positive labels are kept as empty-qrel coverage evidence so the report can distinguish missing qrels from explicitly non-relevant qrels.