Feature/mcts transfer optimizer (Refactored) - #846
Open
jack89roberts wants to merge 13 commits into
Open
jack89roberts wants to merge 13 commits into
jack89roberts wants to merge 13 commits into
Conversation
`squad_for_next_gameweek` sold each outgoing player without `use_api`, so a live run priced the sale from the database while the search had priced it from the API. And it ignored `add_player` returning False: when a sale price came out wrong, a wildcard or free hit's fifteen additions ran out of money partway through, and the squad came back a player or more short. That only surfaced later, as an "incomplete squad" error from lineup selection. It now sells with `use_api` and raises at the addition that fails. From b531763 on mc-tree-search, which fixed the same thing in `print_team_for_next_gw`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
`next_gameweek_transfers`, `count_expected_outputs` and the node step `_make_best_transfers` move from `tree_search.py` to a new `transfer_optimizers/branches.py`, and the node step loses its underscore as `make_best_transfers`. None of them depends on how the tree is walked, and the Monte Carlo search that follows walks the same tree differently. It should reuse the tree's moves and scoring, not import another optimizer's private function. No behaviour changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`--transfer-optimizer mcts` walks the same tree as the exhaustive search, but only the promising parts of it. Each step descends by UCB1 to a node with a move not yet made, makes it with the same strategies and scoring the tree search uses (`branches.make_best_transfers`), and backs up the plan's banked points plus what its squad scores if held to the end of the window. There is no random rollout. Ported from `airsenal/framework/mcts_optimization.py` on mc-tree-search, with four changes that measurement called for: - The result is the best finished plan the search reached, not the most-visited path. Making a move from a node is deterministic, so every finished plan is exactly scored. The most-visited path was worse than the best evaluated plan in both six-gameweek runs (334.24 and 334.39 against 334.53). - The budget is nodes made (`max_expansions`), not iterations. The branch's 300 iterations made only 56 nodes on a three-gameweek window with every chip allowed: most iterations descended to a plan already finished and made nothing. - The search stops once every node is made, rather than spending the rest of its budget revisiting them. - One search, not four. The branch ran four independent copies from the same root with different random orders, which is four times the work for the same tree. It also honours `--max-transfers`, which the branch overrode with the free transfer cap. Measured on the GW6 squad with GW6-11 predictions, against the tree search: | window | tree search | MCTS | |--------------------------|----------------------|---------------------------| | 3 gameweeks, no chips | 185.15, 39 nodes | 185.15, the whole tree | | 6 gameweeks, no chips | 334.53, 1,009 nodes | 334.53 within 100 nodes | | 3 gameweeks, every chip | 199.17, 895 nodes | 198.54 within 56 nodes | | 6 gameweeks, every chip | 163,296 plans | 348.8 within 133 nodes | The first gameweek's move, which is the only one made, was the same in every run. MCTS runs in one process, so on a tree small enough to search exhaustively the four-worker tree search is faster by the clock. Letting every chip be played in any gameweek plays all four within the window: the plan score gives an unplayed chip no value. That is true of either search, so chips should still be pinned to gameweeks. Like any optimizer that is not the default, it starts from its own settings: `--num-thread` and `--num-iterations` reach only the tree search. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
A move of three or more transfers was handled by `random`, which draws `num_iterations` random sets of sales and replacements and keeps the best. On the GW6 squad over GW6-8, from a squad that scores 177.5 if it does nothing: | transfers | random (two seeds) | genetic (two seeds) | |-----------|--------------------|---------------------| | 3 | 180.3 / 176.7 | 183.7 / 181.6 | | 4 | 178.2 / 177.7 | 183.9 / 183.7 | | 5 | 175.3 / 174.8 | 182.2 / 182.2 | Five random transfers do worse than none, so the tree could never usefully take that branch. The genetic search takes about 2.5 s a node, the same as the exhaustive two-transfer search. `SquadOpt` takes a `base_squad` and `max_transfers`. An individual is then the base squad with at most that many players replaced, seeded near it, and with the unchanged squad always in the first population, so there is always a legal answer. Each sale adds its sell price to the bank, and a kept player keeps what was paid for them. The branch this comes from (db427b3 and d8fc765 on mc-tree-search) rebought kept players at a price of zero, which would make every later sale of them look like a profit. An illegal squad scores far below anything legal, rather than zero, so it cannot tie with a legal squad and be returned. `GeneticTransferStrategy` is the new `genetic` entry in `TRANSFER_STRATEGIES`, and `StrategySet.many_transfers` defaults to it. `random` is still registered. The strategy may change fewer players than the move allows, and is charged the full hit when it does; the smaller move is also in the tree and scores better. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
With --max-transfers above 2, the tree branched on every count up to it, and since three or more transfers now each cost a genetic search, a window with five free transfers banked multiplied into far more nodes than it was worth. The tree now branches on 0, 1 and 2 transfers and on using every free transfer. The counts in between are rarely the best use of a bank of transfers. With --max-transfers at its default of 2 the tree is unchanged: those are already every count. It is `TreeSearchConfig.every_transfer_count`, False by default, and `branches.transfer_counts` says which counts it leaves. `count_expected_outputs` takes the same setting, so the progress bar still counts the plans the workers finish. The Monte Carlo search keeps every count, since it only makes the nodes that look promising. From 57950ad on mc-tree-search (`_transfer_count_candidates`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
A search offered each chip once per window. That is right within one half of the season but not across the split: every season gives a wildcard for each half, and from 2025/26 every chip comes twice. A window over gameweeks 19 and 20 never offered the second one. `game/chips.py` holds the rule: `comes_twice(chip, season)`, `season_half`, and `chips_used_up`, which says which of a plan's chips have nothing left to play in a given gameweek. Both searches and `count_expected_outputs` pass that, rather than every chip played so far, and `Plan.chips_by_gameweek` says which gameweek each was played in. From `chip_half` in 621dfdd on mc-tree-search, which applied the split to every chip in every season. Replays of seasons before 2025/26 had only one of each chip other than the wildcard, so the rule depends on the season here. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
A replay runs a fresh search every gameweek, from the same `ChipGameweeks`. Nothing told a search which chips earlier gameweeks had spent, so a replay allowed a wildcard in any gameweek could play one every week: with a stub search that wildcards whenever it may, gameweeks 2 and 3 both did. `ChipGameweeks.played` holds (gameweek, chip) for each chip spent, and the replay adds to it after every gameweek whose plan played one. `TransferSearchRequest.chips_played` carries it into the search, and `TransferSearchRequest.chips_used_up` combines it with the plan so far, so a chip spent before the window is not offered again in the same half of the season. `count_expected_outputs` takes it too. A live run leaves `played` empty and says -1 for a chip it has spent, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`optimization/chip_timing.py` walks the window a gameweek at a time, holding the squad as it is, and decides whether to play a chip in each: - more chips left than gameweeks to play them in before they expire: play one now, in an order that depends on whether the gameweek is blank, double or neither; - a blank that leaves the squad unable to field eleven: free hit, or a wildcard if the blank is severe or the wildcard is about to expire; - a double: bench boost if nearly the whole squad plays twice, triple captain if the captain does with a home game, otherwise a wildcard or free hit depending on what is left and what else is coming. The decided gameweeks are pinned, and the search does the rest. Measuring MCTS showed why this is needed: offered every chip in any gameweek, a search plays all four at once, since a plan's score gives nothing for a chip kept back. `--chip-heuristic` on `optimize transfers`, `run` and `replay` sets `ChipGameweeks.heuristic`, and `run_optimization` decides the gameweeks once the starting squad is known. The chips it starts from are the API's `get_available_chips` for a live run and, for a replay, every chip that `ChipGameweeks.played` has not used up. Each decision is logged with its reason. Ported from `chip_heuristics.py` on mc-tree-search (621dfdd, 099545d, 67b3889, 3bef8e0), with these changes: - A live run started from no record of the chips already played, so it could suggest one that was spent. It now asks the API. - Before 2025/26, only the wildcard expires at gameweek 19; the other chips last until the end of the season. The branch expired every chip at the split in every season. - Fitness is read as at the window's first gameweek, through `Player.is_injured_or_suspended`, so a replay cannot see the future. - A chip decided twice in one window, either side of the split, is pinned to the first gameweek rather than the last. The branch's `--auto_chips`, which offered every available chip in every gameweek, is not ported: with any search, it plays them all at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
A replay wrote its result only once every gameweek had run, so one that failed in gameweek 30 left nothing of the 29 before it. It now rewrites the result after every gameweek, and the final write replaces it. The usual failure was a squad left incomplete, reported as "Squad is incomplete" and nothing more. The error now says how many players the squad has, who they are, and what is in the bank, which is what shows a rebuild that ran out of money. From 8641792, 4306599 and a164763 on mc-tree-search. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: jack89roberts <jroberts@turing.ac.uk>
jack89roberts
force-pushed
the
feature/mcts-transfer-optimizer
branch
from
September 25, 2026 20:50
48da486 to
2212897
Compare
…fer-optimizer # Conflicts: # docs/where-to-look.md
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## develop #846 +/- ##
===========================================
+ Coverage 75.89% 77.34% +1.45%
===========================================
Files 130 136 +6
Lines 7924 8459 +535
Branches 979 1060 +81
===========================================
+ Hits 6014 6543 +529
- Misses 1708 1713 +5
- Partials 202 203 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Tests for what this branch added that nothing ran: - chip_timing: a player who cannot play not counting towards the eleven; a wildcard for a severe or late blank without a free hit; the forced order in a blank or double; and a last chip played on the biggest double before it expires, or kept for a bigger one. - run_optimization: with `--chip-heuristic`, the search is given the chips the heuristic decides, and the chips already played. - The genetic transfer search: a squad player no longer listed is replaced in the same position, and with no transfers allowed there is no legal squad, which is an error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Making a node is nearly all of the search's cost: a strategy run and a scoring, about 3 s, against microseconds to choose and back up. With `MCTSConfig.num_thread` above one (the default is now 4, as for the tree search) this process keeps the tree and a forked `ProcessPoolExecutor` makes the nodes, with up to `num_thread` moves out at once. Two rules keep a search with moves out correct: - A move still out counts as a visit of every node above it when UCB1 chooses between subtrees, so the next choice goes elsewhere rather than down the same path. - A node is not closed off while a move below it is out. Otherwise its subtree is abandoned when the move comes back, and a search can stop with work still unused. One worker makes nodes in this process, as before, and is the only setting a seed makes repeatable: with workers, the order moves come back in is theirs. A worker that dies raises `BrokenProcessPool` instead of hanging. This is not the branch's four root-parallel workers, which each built the same tree. Here there is one tree and the workers share it. Measured on the 2526 replay squad at GW4, six gameweeks, on an otherwise idle M4 Air, against the exhaustive search's optimum of 319.07 over 944 nodes (about 25 minutes): | 4 workers, 200-node budget | nodes | wall clock | |----------------------------|-------|------------| | first move settled | 48 | 77 s | | optimum found | 147 | 200 s | | budget spent | 200 | 5.4 min | Across three squads (GW4, 7, 10) and two seeds, the single-worker search found the exhaustive optimum every time: by node 42-141 of 944 on six gameweeks, and 87-310 of 2,692 on seven. The tests run the workers as threads, which pytest on macOS cannot fork safely. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`--transfer-optimizer auto`, now the default, runs the exhaustive tree search when the tree is small and MCTS once the tree has more than twice as many nodes as the MCTS budget. Both searches make the same nodes at the same cost, so the choice is only how many to make: below twice the budget MCTS would make most of the tree anyway, with no certainty of the best plan. At the default settings (up to 2 transfers, chips pinned) that is the tree search up to five gameweeks (about 330 nodes from one free transfer) and MCTS from six (about 940). The three-gameweek default window is 38 nodes and stays with the tree search. `count_tree_nodes` sizes the tree in nodes, every partial plan, rather than the finished plans `count_expected_outputs` counts: each gameweek's move is a node, and costs about as much to make. Both walk the tree with the same `_tree_levels`. The MCTS budget is 300 by default, up from 200: on seven-gameweek windows the best plan was found within 301-310 nodes on one squad of three. The search flags now reach every optimizer that has them, through its table entry, as `--epsilon` reaches the team models: - `--num-thread` and `--num-iterations`: both searches. - `--max-expansions` (new): the MCTS budget. The tree search rejects it. - `--profile`: the tree search. MCTS rejects it. Under `auto` each flag reaches whichever of the two has it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jack89roberts
marked this pull request as ready for review
September 27, 2026 11:26
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
I asked Claude to transfer the MCTS on top of the refactor branch here @nbarlowATI . Also makes it possible to run in parallel and makes it the default once the size of the whole tree gets beyond a certain size