Skip to content

Latest commit

 

History

History
101 lines (77 loc) · 4.55 KB

File metadata and controls

101 lines (77 loc) · 4.55 KB

Training on SharpeArena

From vf-eval baseline to a GRPO run, on a leak-free trading floor


SharpeArena ships a real, trainable verifiers environment: sharpearena.load_environment() returns an SharpeArenaVerifiersEnv (vf.MultiTurnEnv) that drives the point-in-time market one bar per turn and scores the realized return series with the SharpeBench kernel. This guide covers running it under vf-eval and training on it with prime-rl.

For a copy-pasteable end-to-end config, see examples/prime-rl/.

The env-package layout

The env is a standard installed Python package; the agent contract and Rust engine live behind a pyo3 wheel built by maturin.

crates/sharpearena-py/
  pyproject.toml                     # package metadata + [tool.verifiers.eval] + tags
  python/sharpearena/
    __init__.py
    verifiers_env.py                 # load_environment(), the rubric, the multi-turn rollout
    dataset.py                       # build_scenario_dataset() - the multi-row taskset
    gym.py                           # the underlying gymnasium.Env stepped per turn
    sharpearena_py...pyd              # the compiled Rust kernel (SharpeBench scorer)

load_environment(**args) is the entry point verifiers and prime-rl call. It accepts n_symbols, n_days, n_windows, max_episode_bars, max_turns, max_weight, and allow_short. args.n_windows sets the dataset length, which is the number of GRPO tasks (num_tasks) - it must be > 1 so within-group reward variance has something to vary over.

Install and discover

pip install -e "crates/sharpearena-py[verifiers]"   # editable, with the verifiers extra

Once published to the PrimeIntellect Environments-Hub, the env is installable by org-qualified id:

prime env install general-liquidity/sharpearena

The [project].tags list in pyproject.toml feeds Hub discoverability (the prime CLI reads it at push time); the [tool.verifiers.eval] table sets the default num_examples / rollouts_per_example for a bare vf-eval sharpearena.

Baseline with vf-eval

vf-eval sharpearena \
  -m Qwen/Qwen3-1.7B -n 20 \
  -a '{"n_windows": 20, "n_symbols": 4, "n_days": 120, "max_episode_bars": 16}'

The rubric's dense realized_return_reward is the GRPO objective; the real deflated_sharpe_reward is a secondary objective; pass_k_reward, process_check_reward, and format_reward are zero-weight diagnostics.

v1 taskset and the subprocess runtime

Under the verifiers v1 contract the env is a taskset (taskset = { id = "sharpearena" }) composed with a harness and a runtime. With no container image declared, it runs on the subprocess runtime: prime-rl spawns a local env-server subprocess that imports the installed package and serves rollouts over the worker pool. That is the right default for a pure-Python + pyo3 env with no external services. Drive it from prime-rl with either the v1 taskset = { id = "sharpearena" } form or the legacy id = "sharpearena" form shown in examples/prime-rl/rl.toml.

Leak-freedom: point-in-time is ours, the split is yours

Two distinct guarantees:

  1. Within a scenario - structural, ours. The data layer has no API to read a future bar; the environment owns the time cursor. An agent cannot peek, by construction.
  2. Across train and eval - experimental, yours. A trustworthy training run must evaluate on a strictly held-out seed band. build_scenario_dataset supports this directly: mode="train" draws seeds from base 0, mode="eval" from EVAL_SEED_BASE = 1_000_000, and the two ranges are asserted disjoint (seed_ranges_disjoint).

Caveat (sharpearena 0.1.0). load_environment currently hardcodes mode="train" and does not forward mode/seed_start to build_scenario_dataset. Passing mode = "eval" through a prime-rl eval args block is a silent no-op (the key is absorbed by the verifiers Environment(**kwargs) catch-all), so the eval set reuses the train seed band. Until load_environment forwards mode, treat any in-config eval as not held-out, or build the eval Dataset yourself with build_scenario_dataset(..., mode="eval") and pass it into load_environment(dataset=...) from Python. The eval block in the example config encodes the intended split and is flagged accordingly.