AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
-
Updated
Aug 12, 2026 - Python
AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
Coding Agent Runtime & Evaluation Harness for controlled repository-level software repair
An end-to-end framework for running, sandboxing, and scoring agentic LLMs on complex data-science and econometric replication tasks.
Run 23 speech language models on your own audio task through one interface: vLLM, transformers, and API backends behind a single JSONL contract. Ships HEAR, the speaker-attribution benchmark from our EMNLP 2026 paper.
An sdk and framework for evaluating and comparing multiple model outputs using configurable LLM-based jurors
Teach your agent to work with evals: WHEN you actually need an eval or benchmark, HOW to build one that holds up, and how to read what it tells you. Deterministic-first, tool-agnostic.
Detecting Relational Boundary Erosion in AI systems. A framework for testing whether models maintain honest, calibrated, and appropriate boundaries.
A contract-first PDF extraction and document intelligence evaluation harness.
VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.
Continuous Evaluation Infrastructure for Production AI
Field-level accuracy evaluation of LLM document extraction — POs, packing slips, bills of lading to schema-validated JSON, with committed scoring reports
Autonomous financial research agent combining live market data, financial news, sentiment analysis, and private RAG with transparent execution.
MODA_NER: open fashion attribute extraction — three-track benchmark suite and models by Hopit AI. Companion to hopit-ai/Moda.
Closed-loop LLM factory in one monorepo: data pipeline, trainer, eval-gated checkpoint promotion, serving, and an agent whose single tool is a self-extending CLI.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
Single-file Python library for scanned-document extraction: measured page quality drives adaptive preprocessing, pluggable OCR/VLM backends, schema-driven extraction with bbox provenance, calibrated confidence, and an eval harness. Zero required dependencies.
An LLM agent over crash telemetry that must cite the tool result behind every number it states - with an automatic citation checker and eval harness.
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Wellness verification harness for companion AI. Multi-turn adversarial suites grounded in six decades of mental-health research and current clinical standards (988, VERA-MH) and law (SB 243). Point it at any chat endpoint — get an evidence-backed, reproducible report.
To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."