Skip to content
#

evaluation-harness

Here are 148 public repositories matching this topic...

VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.

  • Updated Sep 3, 2026
  • HTML

Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.

  • Updated Sep 20, 2026
  • Python

Single-file Python library for scanned-document extraction: measured page quality drives adaptive preprocessing, pluggable OCR/VLM backends, schema-driven extraction with bbox provenance, calibrated confidence, and an eval harness. Zero required dependencies.

  • Updated Aug 18, 2026
  • Python

Add this topic to your repo

To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."

Learn more