Mettle is an on-chain reputation layer for AI trading agents. Five autonomous agents each run their own strategy on Mantle, and every move they make is settled and scored on-chain — together with the reasoning behind it. Over time each agent earns a transparent 0–100 track record that anyone can verify, and capital is routed toward the agents that have actually earned it.
Mettle in one sentence: instead of asking you to trust an agent's marketing, Mettle makes the agent prove itself, in public, trade by trade.
"AI trading agent" is one of the easiest things in the world to fake. Anyone can post a screenshot of a green P&L curve, claim an impressive win rate, or wrap a coin-flip in confident language. There is usually no way to tell a genuinely good strategy from a lucky one, or from an outright lie — and by the time you can, your money is already in.
Mettle removes the need to trust any of that. An agent doesn't get to tell you it's good; it has to demonstrate it where no one can edit the result:
- Its trades happen on-chain, inside a vault it cannot withdraw from.
- Its score is computed on-chain from realized profit and loss, not self-reported.
- Its reasoning for every decision is fingerprinted on-chain, so it can't be rewritten after the outcome is known.
What you're left with is a reputation you can audit yourself.
Each agent owns one vault. A scoring round (an "epoch") runs like this:
- Read the market. An off-chain service pulls real recent prices for each tradable asset.
- Decide. A language model, prompted to act as that agent's strategy, picks one asset to go long for the round and a size — or chooses to sit in cash. It returns a short rationale in its own voice.
- Risk-check. Off-chain risk limits cap the size, reject low-conviction or malformed calls, and force cash when nothing fits. Nothing reaches the chain unchecked.
- Settle on-chain. The vault opens an epoch, runs the trade through a simple on-chain market, lets the real market move play out, and closes back to cash.
- Score. The vault measures its own realized USD profit and loss for the round and maps it to a 0–100 score (50 is breakeven). It writes that score to the validation registry.
- Build reputation. Scores accumulate into a track record. A separate allocation controller then routes pooled capital toward the agents whose records clear a minimum bar — and away from the underperforming ones.
The agent's rationale is hashed and the hash stored alongside the score, so the words it used to justify a trade are locked in before anyone knows whether the trade worked.
You don't have to pick an agent. The AllocationController is a pooled index: deposit USD once and you get index shares representing a claim on the whole book. After each round the operator rebalances the index automatically — it recalls the deployed capital and re-deploys it across the eligible agents weighted by their freshly-updated scores, so money continuously flows toward whoever is proving themselves and away from those who aren't. An agent below the eligibility bar (a minimum track record and an average score at or above breakeven) draws nothing.
Because the rebalance runs on the same schedule as the rounds, "deposit into the index" genuinely means "let the on-chain track record decide where my money goes" — no manual allocation, no trusted manager. You can still deposit directly into a single agent's vault if you want to back one strategy; the index is simply the hands-off option. Withdrawals are paid from the index's idle capital, which is replenished each rebalance.
Mettle is built on ERC-8004, the standard for on-chain agent identity, reputation, and validation. ERC-8004 gives agents an identity (an NFT) and a place to record validation results, but it leaves open the hard question: who validates a trading agent honestly?
Mettle's answer is that the vault validates itself, and it can't cheat while doing so. Between epochs the vault holds nothing but USD, so its starting and ending balances are unambiguous. The difference between them is realized profit and loss — there is no oracle to trick and no subjective judgement to game. The vault opens its own ERC-8004 validation request at the start of an epoch and answers it with the measured score at the end. The score is therefore as trustworthy as arithmetic.
This is what makes the reputation meaningful rather than decorative: it is derived, on-chain, from money that genuinely moved.
Five agents ship in the seed deployment, each mapped to a Mantle-native asset:
| Agent | Strategy | Asset |
|---|---|---|
| Momentum Alpha | Rides strong trends | mETH |
| Breakout Hunter | Buys breakouts from ranges | fBTC |
| Volatility Harvester | Trades large swings | MNT |
| Steady Yield | Capital preservation | USDY |
| Mean Reversion | Fades overextended moves | MI4 (a Mantle index) |
On testnet the assets are mock ERC-20s priced through Mettle's own market contract, so rounds are deterministic and self-contained. The off-chain brain still reasons over real recent price action for the corresponding assets (ETH, BTC, MNT, and a BTC/ETH/SOL blend for the index), sourced from Bybit with a CoinGecko fallback.
- IdentityRegistry — ERC-721 identities for agents (the ERC-8004 identity layer).
- ReputationRegistry — append-only feedback records keyed to an agent.
- ValidationRegistry — validation requests and responses; holds each epoch's score and can summarize an agent's record filtered to a chosen validator.
- StrategyVault — a non-custodial, single-agent vault that trades, measures its own realized P&L, and writes its score. The heart of the system.
- VaultFactory — launches an agent (mints its identity and deploys its official vault) and tracks which vaults are official.
- Market — a minimal on-chain venue that swaps between USD and the tradable tokens at a settable price.
- AIRunner — the on-chain operator that takes an AI decision (asset, size, real move, rationale URI and hash), drives one epoch through a vault, and logs the decision.
- AllocationController — a pooled USD index that routes capital into official vaults weighted by validation score, behind track-record and quality gates.
- agent/ — the off-chain service: market data, the model-driven decision brain, the risk layer, and the round runner that executes decisions on-chain.
src/
IdentityRegistry.sol ERC-721 agent identities (ERC-8004 identity)
ReputationRegistry.sol append-only feedback per agent
ValidationRegistry.sol validation requests/responses; holds scores
StrategyVault.sol non-custodial vault; the "vault is the validator"
VaultFactory.sol launches agents and their official vaults
Market.sol minimal USD <-> token swap venue
AIRunner.sol runs one AI decision through a vault, on-chain
AllocationController.sol routes pooled capital by validation score
interfaces/ registry and market interfaces
mocks/MockERC20.sol testnet tokens
script/
Deploy.s.sol full-stack deploy + seeds five agents
test/ Foundry tests for vaults, registries, runner, allocation
agent/
src/config.ts chain, addresses, assets, strategy personas, risk limits
src/market.ts real price feeds (Bybit primary, CoinGecko fallback)
src/brain.ts the model-driven decision per agent
src/risk.ts off-chain risk checks before anything goes on-chain
src/run.ts one full round: decide, check, execute, record
deployed.json live Mantle Sepolia addresses
forge build
forge testDeploy the full stack and seed the five agents:
forge script script/Deploy.s.sol --rpc-url mantle_sepolia --account deployer --broadcast --verifycd agent
npm install
cp .env.example .env # fill in the LLM endpoint + the operator key
npm run roundnpm run round reads the market, asks the model for each agent's move, runs the risk checks, executes each decision on-chain, and writes a full record of the round (decisions, rationales, scores, transaction hashes) to agent/rationales/. After the round it rebalances the pooled index (see below).
To run rounds autonomously on a timer instead of by hand, use loop mode:
npm run loop # a round every 30 minutes
ROUND_INTERVAL_MINUTES=120 npm run round # or pick your own intervalThe loop runs a round, rebalances the index, sleeps, and repeats until stopped. A failed round is logged but doesn't kill the loop — the next tick tries again.
Auto-allocation. After each round the runner recalls the index's deployed capital and re-deploys it across the eligible agents weighted by their freshly-settled scores, so a plain deposit into the index flows to the best performers with no manual step. allocate is owner-gated, so the AllocationController's ownership must be transferred to the operator for this to run; until then the runner logs that it's skipping allocation and the rounds proceed normally.
The runner uses a dedicated, low-risk operator key. The operator controls only the AIRunner (and, once ownership is transferred, the AllocationController), both of which are non-custodial — they can trade inside vaults and route pooled capital, but can never move anyone's funds out — so this key is deliberately separate from the deployer.
Keeping the operator funded. Each round is one transaction per agent plus a recall and an allocate. L2 execution is only ~640k gas per call, but Mantle also bills an L1 data fee for the calldata — the rationale string dominates it — so a measured round settles at about 3.2 MNT. At three rounds a day that is roughly 10 MNT a day, and the operator needs periodic top-ups from the Mantle Sepolia faucet. The runner checks the balance before every round and refuses to start with a clear message when it can't cover one, and warns while under three days of runway remains. This matters because of how the failure looks otherwise: the node caps eth_estimateGas at balance / gasPrice, the call runs out of gas at that cap, and the RPC reports it as execution reverted — which surfaces as "the contract function runEpochAI reverted" and looks like a contract bug rather than an empty wallet.
Explorer: https://sepolia.mantlescan.xyz
Core contracts:
| Contract | Address |
|---|---|
| Market | 0xba61bbdc03c3df64e256186c5187e52b09262dc2 |
| IdentityRegistry | 0xd279d843ccc1908bbf1f470fe37e2b22155300b1 |
| ReputationRegistry | 0xcd474c41a48ffa6b6296899827f8b274e1c0a56d |
| ValidationRegistry | 0xd4296d8ced0644fa29615e3d342853ee955e696a |
| VaultFactory | 0x6c5f6f0e683dad2b318b78d0eb1bef816f55d895 |
| AIRunner | 0xb3b1a270be197a46ab2c63c41e700fdb07be7f6e |
| AllocationController | 0x5843d11bb0d95cb16ce1ba1fa9448ffac5fcbef5 |
Agent vaults:
| Agent | Vault |
|---|---|
| Momentum Alpha | 0x5fe4cdd6c12712968cb90a6e513417d55c0f8cdd |
| Breakout Hunter | 0x0f3e55fd68a17ad653f51f810728b0c8a60cdf8f |
| Volatility Harvester | 0x1d665641a18ed29efd6377af56f4510f3f53cd31 |
| Steady Yield | 0xda2392671d08e7f15cad73697ff54cd03755a02b |
| Mean Reversion | 0x3ea332055fef9545191bff1a11f7eac20cb2141b |
All token, core, and vault addresses are kept in deployed.json.
A reputation system is only as good as the things it refuses to be fooled by. The vault is built to be donation-proof and self-contained:
- Scores come from accounted P&L, never
balanceOf. Someone sending USD straight to a vault cannot inflate its score or its share price. - Capital is ring-fenced per epoch. Only the capital the vault accounts for can be traded; donated tokens can't enter the trade flow.
- The agent can trade your money but never take it. The vault has no path to send funds to an arbitrary address — tokens only move vault → market → vault.
- Allocation reads filtered scores. The controller counts only a vault's own self-validations and gates on a minimum track record and score, so one lucky epoch or a fabricated score can't draw capital.
- Allocation-driven competition where capital continuously chases the best live records.
- Richer strategy personas and multi-asset positions per round.
- A swap to live on-chain venues and real assets beyond the testnet mocks.
- A public transparency dashboard for watching agents decide, score, and earn capital in real time.
MIT.