Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

altiquum-eval

A reproducible evaluation framework for AI systems, retrieval systems, agents, and model experiments.

Why

Agent workflows require standardized, reproducible testing to prevent behavioral regressions and verify model accuracy.

Goals

  • Provide structured model evaluation templates\n- Track dataset test metrics over iteration runs\n- Implement pipeline assertion checks\n

Non-Goals

  • nongoal\n- nongoal\n

Architecture

A Python framework using pytest structures and JSON schemas to define, evaluate, and track run benchmarks.

Getting Started

Clone the repository and see language-specific tool setups.

Usage

Refer to examples directory for minimal usage.

Development

See CONTRIBUTING.md for details.

Testing

Run language-specific testing framework commands.

Roadmap

  • Establish baseline interfaces
  • Add unit test coverage
  • Integrate telemetry tracking

Contributing

See CONTRIBUTING.md.

License

Apache-2.0 License.

About

A reproducible evaluation framework for AI systems, retrieval systems, agents, and model experiments.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors