OEV is a small (184M params) neural decision engine. Instead of generating text, it scores answer options directly. The state, the question, and every option are packed into one sequence. One forward pass returns a calibrated probability distribution.
The single 184M model scores 0.7705 on typed-decisions, slightly above laya's published 0.766 from a 421M checkpoint. An ensemble of four 184M checkpoints reaches 0.7760, the highest reported result, and a single Banking77 soup checkpoint reaches 0.8594 (best ECE 0.0583 from the 4-checkpoint ensemble). Jev leads only on Banking77 (0.870).
| claim | result |
|---|---|
| best accuracy | 0.7705 single model, 0.7760 ensemble - typed-decisions (laya 0.766 from 421M) |
| high-cardinality | 0.8594 Banking77, single 184M file (ECE 0.0583 best ensemble; laya 0.425) |
| speed | 22.2 ms single question (laya 32.8-39.5 ms published range) |
| zero-shot emotion | 0.8650 DAIR Emotion zero-shot (laya 0.595) |
| size | 184M params, 0.44x laya |
| weights & checkpoints | Apache 2.0 - huggingface.co/divyanshudhruv/oev-typed |
22.2 msper question on aT4(GPU); on CPU the ONNX INT8 build runs at54.2 msp50 on 8 threads (228 MBartifact, 3.2x smaller)0.8594on 77-labelBanking77from a single file (the ensemble holds best ECE0.0583): each option is embedded as its own anchor with full tokens, so accuracy scales with label count (gap to Jev 1.06 pts)184Mparams,Apache 2.0weights- Kev (0.8B / 4B) publishes no in-domain numbers on these datasets, so it is not in the tables; see BENCHMARKS.md for the like-for-like comparison plan
flowchart LR
subgraph input["Input"]
direction TB
S["state<br/>text or JSON"]
Q["questions<br/>choice, noul, score"]
end
P["packer<br/>state + questions + anchors<br/>one packed sequence"]
subgraph pass["One forward pass, 22 ms on T4"]
direction TB
E["encoder<br/>DeBERTa-v3-base, 184M"]
H["shared linear head<br/>scores every anchor"]
end
D["softmax per question<br/>calibrated distribution"]
O["outputs<br/>choice / noul / score"]
S --> P
Q --> P
P --> E
E --> H
H --> D
D --> O
- One anchor mechanism covers all three primitives - options, yes/no pairs, and score levels are each embedded as anchors in one packed sequence
- No text generation - nothing to parse, nothing to hallucinate
- New question types need no new heads - new options are just new anchors
A from-scratch character-level encoder also exists for the no-transformers path (see oev/tokenizer.py); the diagrams and benchmarks here all use the DeBERTa backbone.
A generative model answers a structured question like this:
input -> generate text -> parse the output -> validate it -> maybe get a decision
OEV is built for the case where the system already knows the candidate answers:
state + candidates -> score every candidate -> calibrated probability distribution
If you need open-ended text, use an LLM. If you need many small, bounded decisions - which department, is this refund requested, how severe, escalate or continue - scoring known options directly is cheaper, faster, and structurally immune to output-parsing failures: the outputs are probabilities over the inputs you supplied. Different tool for a different layer of the stack.
Fine-tuned on each benchmark's train split, following the same protocol as Laya's published runs. Selected results and evaluation notes are in BENCHMARKS.md.
Warning
Banking77 uses 77 OEV labels, while the published Jev figure is from a 72-label configuration. The 0.8594 and 0.870 values are not a controlled head-to-head comparison.
| benchmark | OEV | laya | Jev | note |
|---|---|---|---|---|
| typed-decisions | 0.7760 | 0.766 | 0.727 | highest reported (ensemble); single model 0.7705 |
| Banking77 | 0.8594 | 0.425 | 0.870 | 2x laya; gap to Jev 1.06 pts |
| AG News | 0.9489 | 0.950 | 0.910 | label-noise ceiling (~0.95) |
| DAIR Emotion | 0.9300 (fine-tuned) | 0.595 | 0.480 | zero-shot: OEV round-3 student 0.8650 beats laya's 0.595 head-to-head |
pip install oev # from PyPI
# or from source:
pip install -e .Optional extras:
pip install -e ".[dev]"- pytestpip install -e ".[data]"- dataset converterspip install -e ".[backbone]"- DeBERTa fine-tuningpip install -e ".[serve]"- FastAPI server (oev-serve)pip install -e ".[app]"- Gradio Space dependencies
from oev import OEV
agent = OEV("checkpoints_td5/oev-tiny.pt", device="cpu")
result = agent.decide("We were charged twice for the same order.", {
"department": {"type": "choice", "options": ["billing", "technical", "sales", "other"],
"instructions": "Which department should handle this?"},
"refund_requested": {"type": "noul"},
"severity": {"type": "score", "levels": [1, 2, 3, 4, 5]},
})
# Native schema above; the Jev/TypeSafe `criteria` schema (the one laya and Kev
# speak) works unchanged - drop-in for existing Jev clients:
result = agent.decide("We were charged twice for the same order.", {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments", "technical": "bugs, outages"}},
"refund_requested": {"type": "noul"},
"severity": {"type": "score", "criteria": ["minor", "soon", "blocking"]},
})
# The candidate set is an inference-time input - swap it per call, same weights:
state = "play my morning playlist and take a note"
actions = ["open_spotify", "search_web", "open_vscode", "create_note"]
result = agent.decide(state, {"action": {"type": "choice", "options": actions}}){
"department": {
"choice": "billing",
"probabilities": {
"billing": 0.94,
"technical": 0.04,
"sales": 0.01,
"other": 0.01
},
"confidence": 0.94
},
"refund_requested": 0.91,
"severity": {
"value": 3,
"probabilities": { "1": 0.02, "2": 0.08, "3": 0.61, "4": 0.22, "5": 0.07 }
}
}Confidence gating - automate when confident, escalate when not:
from oev.presets import triage_questions, gate
for name, payload, confident in gate(result, threshold=0.85):
automate(name, payload) if confident else escalate_to_human(name)HTTP server (native + Jev-compatible /v1/systemone endpoint - TypeSafe clients work by changing baseUrl):
pip install -e ".[serve]"
oev-serve --checkpoint checkpoints_td5/oev-tiny.pt --port 8000
curl -X POST localhost:8000/decide -H "Content-Type: application/json" \
-d '{"state": "My payment failed twice", "questions": {"urgency": {"type": "score", "criteria": ["not urgent", "soon", "blocking"]}}}'Docker:
docker compose up # checkpoint at ./checkpoints/oev-tiny.ptpython -m oev.convert_typed
python -m oev.train --backbone microsoft/deberta-v3-base --epochs 4 --batch-size 8 --max-len 768 --data-dir data/typed --out checkpoints_td5
python -m oev.rlcd --checkpoint checkpoints_td5/oev-tiny.pt --data-dir data/typed --epochs 2 --batch-size 8 --out checkpoints_rlcd
python -m oev.ensemble --ckpts checkpoints_td5/oev-tiny.pt,checkpoints_rlcd/oev-tiny.pt,checkpoints_rlcd_soup/oev-tiny.pt,checkpoints_rlcd_seed1/oev-tiny.pt --data-dir data/typed
python -m pytest -qtrain_colab.ipynb runs the entire pipeline end to end. Evaluation rules and scope notes are in BENCHMARKS.md.
- Push Banking77 past Jev's
0.870: current best0.8594- lineage diversity, data augmentation or distillation, not more warm-starts - Multi-question shared-state encoding (one pass, many questions)
- Robustness: reduce overconfidence on out-of-distribution inputs
- Non-English checkpoints (the interface is language-agnostic; the weights are not yet)
Full list: ROADMAP.md.
- Benchmarks - methodology, caveats, reproduction commands, checkpoint hashes
- Model card - checkpoints and usage
- Roadmap - what is next
- Contributing - how to open issues and PRs
- Changelog - release history
- Security - supported versions and private reporting
- Hugging Face demo - try it in the browser
- PyPI -
pip install oev - Server API -
oev-serve --help, schema docs inoev/serve.py
The interface and benchmark protocol follow Laya, Kev and the System One model category introduced by TypeSafe's Jev. Their published numbers are quoted here for comparison and remain their measurements.
