Skip to content
divyanshudhruvPublic

About

An 184M-parameter decision engine delivering calibrated distributions from one 22ms forward pass while matching models 2.3x its size.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

84 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

OEV

OEV is a small (184M params) neural decision engine. Instead of generating text, it scores answer options directly. The state, the question, and every option are packed into one sequence. One forward pass returns a calibrated probability distribution.

The single 184M model scores 0.7705 on typed-decisions, slightly above laya's published 0.766 from a 421M checkpoint. An ensemble of four 184M checkpoints reaches 0.7760, the highest reported result, and a single Banking77 soup checkpoint reaches 0.8594 (best ECE 0.0583 from the 4-checkpoint ensemble). Jev leads only on Banking77 (0.870).

Hugging Face Model License Tests CodeQL Python PyTorch HF Space PyPI

OEV vs Jev and laya on shared public benchmarks

OEV versus TypeSafe Jev: accuracy on shared public datasets, every application workflow, speed, calibration, size, and soft-accuracy sharpening

Left: zero-shot and out-of-domain transfer, OEV distilled student beats laya zero-shot on emotion with the NLI floor documented; right: latency, OEV 22.2ms on T4 vs laya, Kev-4B and Jev

At a glance

claim result
best accuracy 0.7705 single model, 0.7760 ensemble - typed-decisions (laya 0.766 from 421M)
high-cardinality 0.8594 Banking77, single 184M file (ECE 0.0583 best ensemble; laya 0.425)
speed 22.2 ms single question (laya 32.8-39.5 ms published range)
zero-shot emotion 0.8650 DAIR Emotion zero-shot (laya 0.595)
size 184M params, 0.44x laya
weights & checkpoints Apache 2.0 - huggingface.co/divyanshudhruv/oev-typed
  • 22.2 ms per question on a T4 (GPU); on CPU the ONNX INT8 build runs at 54.2 ms p50 on 8 threads (228 MB artifact, 3.2x smaller)
  • 0.8594 on 77-label Banking77 from a single file (the ensemble holds best ECE 0.0583): each option is embedded as its own anchor with full tokens, so accuracy scales with label count (gap to Jev 1.06 pts)
  • 184M params, Apache 2.0 weights
  • Kev (0.8B / 4B) publishes no in-domain numbers on these datasets, so it is not in the tables; see BENCHMARKS.md for the like-for-like comparison plan

Architecture

flowchart LR
    subgraph input["Input"]
        direction TB
        S["state<br/>text or JSON"]
        Q["questions<br/>choice, noul, score"]
    end

    P["packer<br/>state + questions + anchors<br/>one packed sequence"]

    subgraph pass["One forward pass, 22 ms on T4"]
        direction TB
        E["encoder<br/>DeBERTa-v3-base, 184M"]
        H["shared linear head<br/>scores every anchor"]
    end

    D["softmax per question<br/>calibrated distribution"]
    O["outputs<br/>choice / noul / score"]

    S --> P
    Q --> P
    P --> E
    E --> H
    H --> D
    D --> O
Loading
  • One anchor mechanism covers all three primitives - options, yes/no pairs, and score levels are each embedded as anchors in one packed sequence
  • No text generation - nothing to parse, nothing to hallucinate
  • New question types need no new heads - new options are just new anchors

A from-scratch character-level encoder also exists for the no-transformers path (see oev/tokenizer.py); the diagrams and benchmarks here all use the DeBERTa backbone.

Why not an LLM?

A generative model answers a structured question like this:

input -> generate text -> parse the output -> validate it -> maybe get a decision

OEV is built for the case where the system already knows the candidate answers:

state + candidates -> score every candidate -> calibrated probability distribution

If you need open-ended text, use an LLM. If you need many small, bounded decisions - which department, is this refund requested, how severe, escalate or continue - scoring known options directly is cheaper, faster, and structurally immune to output-parsing failures: the outputs are probabilities over the inputs you supplied. Different tool for a different layer of the stack.

Benchmarks: OEV vs the published field

Fine-tuned on each benchmark's train split, following the same protocol as Laya's published runs. Selected results and evaluation notes are in BENCHMARKS.md.

Warning

Banking77 uses 77 OEV labels, while the published Jev figure is from a 72-label configuration. The 0.8594 and 0.870 values are not a controlled head-to-head comparison.

benchmark OEV laya Jev note
typed-decisions 0.7760 0.766 0.727 highest reported (ensemble); single model 0.7705
Banking77 0.8594 0.425 0.870 2x laya; gap to Jev 1.06 pts
AG News 0.9489 0.950 0.910 label-noise ceiling (~0.95)
DAIR Emotion 0.9300 (fine-tuned) 0.595 0.480 zero-shot: OEV round-3 student 0.8650 beats laya's 0.595 head-to-head

OEV headline results: typed-decisions accuracy, Banking77 accuracy, and hardware-separated latency

Zero-shot and out-of-domain transfer as dot pairs: OEV 0.650 versus laya 0.595 on emotion, with the WANLI and ANLI floors marked OEV latency on Tesla T4 and CPU, hardware separated; batch value is per-question throughput

Typed-decisions accuracy per workflow: OEV leads three of four workflows, including invoice processing and customer service Illustrative normalized distributions for OEV choice, noul, and score decision primitives

Quickstart

pip install oev          # from PyPI
# or from source:
pip install -e .

Optional extras:

  • pip install -e ".[dev]" - pytest
  • pip install -e ".[data]" - dataset converters
  • pip install -e ".[backbone]" - DeBERTa fine-tuning
  • pip install -e ".[serve]" - FastAPI server (oev-serve)
  • pip install -e ".[app]" - Gradio Space dependencies
from oev import OEV

agent = OEV("checkpoints_td5/oev-tiny.pt", device="cpu")

result = agent.decide("We were charged twice for the same order.", {
    "department": {"type": "choice", "options": ["billing", "technical", "sales", "other"],
                   "instructions": "Which department should handle this?"},
    "refund_requested": {"type": "noul"},
    "severity": {"type": "score", "levels": [1, 2, 3, 4, 5]},
})

# Native schema above; the Jev/TypeSafe `criteria` schema (the one laya and Kev
# speak) works unchanged - drop-in for existing Jev clients:
result = agent.decide("We were charged twice for the same order.", {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments", "technical": "bugs, outages"}},
    "refund_requested": {"type": "noul"},
    "severity": {"type": "score", "criteria": ["minor", "soon", "blocking"]},
})

# The candidate set is an inference-time input - swap it per call, same weights:
state = "play my morning playlist and take a note"
actions = ["open_spotify", "search_web", "open_vscode", "create_note"]
result = agent.decide(state, {"action": {"type": "choice", "options": actions}})
{
  "department": {
    "choice": "billing",
    "probabilities": {
      "billing": 0.94,
      "technical": 0.04,
      "sales": 0.01,
      "other": 0.01
    },
    "confidence": 0.94
  },
  "refund_requested": 0.91,
  "severity": {
    "value": 3,
    "probabilities": { "1": 0.02, "2": 0.08, "3": 0.61, "4": 0.22, "5": 0.07 }
  }
}

Confidence gating - automate when confident, escalate when not:

from oev.presets import triage_questions, gate

for name, payload, confident in gate(result, threshold=0.85):
    automate(name, payload) if confident else escalate_to_human(name)

HTTP server (native + Jev-compatible /v1/systemone endpoint - TypeSafe clients work by changing baseUrl):

pip install -e ".[serve]"
oev-serve --checkpoint checkpoints_td5/oev-tiny.pt --port 8000
curl -X POST localhost:8000/decide -H "Content-Type: application/json" \
  -d '{"state": "My payment failed twice", "questions": {"urgency": {"type": "score", "criteria": ["not urgent", "soon", "blocking"]}}}'

Docker:

docker compose up   # checkpoint at ./checkpoints/oev-tiny.pt

Training

python -m oev.convert_typed

python -m oev.train --backbone microsoft/deberta-v3-base --epochs 4 --batch-size 8 --max-len 768 --data-dir data/typed --out checkpoints_td5

python -m oev.rlcd --checkpoint checkpoints_td5/oev-tiny.pt --data-dir data/typed --epochs 2 --batch-size 8 --out checkpoints_rlcd

python -m oev.ensemble --ckpts checkpoints_td5/oev-tiny.pt,checkpoints_rlcd/oev-tiny.pt,checkpoints_rlcd_soup/oev-tiny.pt,checkpoints_rlcd_seed1/oev-tiny.pt --data-dir data/typed

python -m pytest -q

train_colab.ipynb runs the entire pipeline end to end. Evaluation rules and scope notes are in BENCHMARKS.md.

Roadmap

  • Push Banking77 past Jev's 0.870: current best 0.8594 - lineage diversity, data augmentation or distillation, not more warm-starts
  • Multi-question shared-state encoding (one pass, many questions)
  • Robustness: reduce overconfidence on out-of-distribution inputs
  • Non-English checkpoints (the interface is language-agnostic; the weights are not yet)

Full list: ROADMAP.md.

Documentation

Credits

The interface and benchmark protocol follow Laya, Kev and the System One model category introduced by TypeSafe's Jev. Their published numbers are quoted here for comparison and remain their measurements.


OEV banner

About

An 184M-parameter decision engine delivering calibrated distributions from one 22ms forward pass while matching models 2.3x its size.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages