Falsiflow does not need to run your local model, expose an API port, or receive cloud credentials. Keep Ollama, LM Studio, llama.cpp, MLX, vLLM, or an internal runner on your side of the boundary. Falsiflow consumes the evidence artifacts that runner already produced.
Use this when a claim such as "the local model improved" should stay blocked in CI until the repository provides reviewable provenance, raw outputs, metric rows, and reproducibility metadata.
A local LLM eval handoff should include:
- dataset or task-set version;
- prompt-set hash;
- candidate model id and local model file hash;
- baseline model id or revision;
- evaluator or harness version;
- raw output artifact path;
- item count, score, hallucination or safety boundary, and run metadata;
- deterministic decode settings, seed, or an explicit note when the run is not deterministic.
Those values can come from a JSON manifest, JSONL metric rows, CSV exports, or a small wrapper script around your local runner.
pipx install falsiflow
falsiflow quickstart --template ai_claim_evaluation --out local_llm_eval_gate --strict
cd local_llm_eval_gateStore the raw eval export under source_files/ so the source manifest can prove
the evidence file exists:
cat > source_files/local_eval_results.jsonl <<'JSONL'
{"model_id":"candidate_model","metric":"exact_match_rate","value":0.86}
{"model_id":"candidate_model","metric":"hallucination_rate","value":0.035}
{"model_id":"candidate_model","metric":"safety_policy_failure_rate","value":0.012}
{"model_id":"candidate_model","metric":"evaluated_item_count","value":640}
{"model_id":"baseline_model","metric":"exact_match_rate","value":0.78}
{"model_id":"baseline_model","metric":"hallucination_rate","value":0.07}
{"model_id":"baseline_model","metric":"safety_policy_failure_rate","value":0.025}
{"model_id":"baseline_model","metric":"evaluated_item_count","value":640}
JSONLAdd the run manifest from your local runner:
cat > local_model_manifest.json <<'JSON'
{
"eval_run_id": "eval_run_001",
"candidate_model_id": "candidate_model",
"baseline_model_id": "baseline_model",
"dataset_version": "claims_eval_v2026_05_26",
"prompt_set_hash": "promptset_sha256_demo",
"model_file_hash": "gguf_sha256_demo",
"baseline_model_version": "baseline_llm_2026_05_01",
"evaluator_version": "eval_harness_0.4.0",
"eval_script_hash": "eval_script_sha256_demo",
"decode_params": {"temperature": 0, "top_p": 1, "seed": 7},
"raw_outputs_uri": "source_files/local_eval_results.jsonl",
"human_spotcheck_passed": true,
"ci_run_url": "local://manual-run",
"runtime": "llama.cpp",
"quantization": "Q4_K_M",
"measured_at": "2026-05-26T12:00:00Z"
}
JSONImport the local runner artifacts into the Falsiflow evidence contract:
falsiflow evidence import \
--profile local-llm-eval \
--input source_files/local_eval_results.jsonl \
--manifest local_model_manifest.json \
--out evidence.csv \
--config project.json \
--coverage-out import_coverage.json \
--source-file source_files/local_eval_results.jsonl \
--strictThen run the same checks a reviewer or CI gate can run:
falsiflow doctor --project-dir . --strict
falsiflow claim-check --config project.json --evidence evidence.csv --out-dir claim_check --strict --forceExpected result:
doctor_ready
claim_check_ready
bundle_verified
Common local runner fields map naturally into the Falsiflow manifest:
| Runner fact | Manifest field |
|---|---|
| Ollama, LM Studio, llama.cpp, MLX, vLLM, or internal harness name | runtime |
| GGUF, safetensors, adapter, or model artifact hash | model_file_hash or adapter_hash |
| Quantization method | quantization |
| Temperature, top-p, max tokens, seed | decode_params |
| Eval dataset, task set, or prompt collection | dataset_version, prompt_set_hash |
| Harness or judge revision | evaluator_version, eval_script_hash |
| Raw model outputs | raw_outputs_uri |
If your local runner emits a different JSON shape, keep a tiny normalization
script outside Falsiflow that writes local_eval_results.jsonl and
local_model_manifest.json. Falsiflow should remain the evidence gate, not the
model runner.
In CI, commit or upload sanitized eval artifacts, then run the reusable action:
- uses: AzurLiu/falsiflow@v0.1.16
with:
mode: claim-check
project-dir: local_llm_eval_gate
evidence: local_llm_eval_gate/evidence.csv
strict: "true"claim_check_ready means the configured evidence package is complete enough for
review. It does not prove the local model is safe, correct, non-hallucinating,
or better for your product.