-
Notifications
You must be signed in to change notification settings - Fork 80
All issues
Issue creation is restricted in this repository
- #597 · utilForever opened
on Jun 10, 2021
Issues
is:issue state:open
is:issue state:open
Search results
[RL][Phase 6] Validate and document the complete AlphaZero MVP
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1084 In utilForever/RosettaStone;[RL][Phase 6] Promote models safely and record reproducible run metadata
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1083 In utilForever/RosettaStone;[RL][Phase 6] Implement paired-seat arena evaluation
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1082 In utilForever/RosettaStone;[RL][Phase 6] Orchestrate one AlphaZero training iteration
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1081 In utilForever/RosettaStone;[RL][Phase 5] Implement bounded replay and an end-to-end self-play smoke test
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1080 In utilForever/RosettaStone;[RL][Phase 5] Add AlphaZero root exploration controls
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1079 In utilForever/RosettaStone;[RL][Phase 5] Implement the fixed-deck mirror-match episode runner
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1078 In utilForever/RosettaStone;[RL][Phase 5] Define the self-play sample format
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1077 In utilForever/RosettaStone;[RL][Phase 4] Verify C++ and Python training parity
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1076 In utilForever/RosettaStone;[RL][Phase 4] Bind the C++ model and implement the Python trainer
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1075 In utilForever/RosettaStone;[RL][Phase 4] Implement the native C++ trainer
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1074 In utilForever/RosettaStone;[RL][Phase 4] Define a cross-runtime model-weight artifact
C-rlCategory: Search, self-play, reinforcement learning, and training.Category: Search, self-play, reinforcement learning, and training.P-importantPriority: Other work depends on this, or it is low-level and critical.Priority: Other work depends on this, or it is low-level and critical.T-featureType: New capability or supported behavior.Type: New capability or supported behavior.Status: Open.#1073 In utilForever/RosettaStone;