|
1 | | -[](https://classroom.github.com/a/64wEcMIk) |
2 | | -# <img src="frontend/favicon.svg" alt="logo" width="128" height="128" align="middle"> SI2 - Othello |
| 1 | +# Intelligent Systems II: Autonomous Othello Agents |
3 | 2 |
|
4 | | -Othello is a classic strategy board game played on an 8x8 grid. The game involves two players, Black and White, who take turns placing their discs on the board. The primary objective is to out-position your opponent by "sandwiching" their pieces between your own, which allows you to flip them to your color. The player with the most discs of their color on the board when no more moves can be made wins the game. |
| 3 | +**Authors:** Francisco Carvalho (114492) & João Viegas (113144) |
5 | 4 |
|
6 | | -This project provides a simulation environment for Othello, featuring a backend server that manages the game state, a frontend for visualization, and a framework for developing autonomous agents. The backend handles the game logic, enforces rules, and communicates with connected agents via WebSockets, while the frontend provides a real-time view of the board and match statistics. |
| 5 | +## 1. Project Overview & Execution Instructions |
7 | 6 |
|
8 | | -## Game Rules |
| 7 | +This project implements an autonomous agent for the classic board game Othello. Our final submission utilizes an **Efficiently Updatable Neural Network (NNUE)** trained via Deep Q-Learning (DQN). |
9 | 8 |
|
10 | | -The game is played on an 8x8 board. It starts with four discs placed in a square in the center of the grid: two white and two black. Black always moves first. |
| 9 | +To run our agents against the simulation server, follow these instructions: |
11 | 10 |
|
12 | | -- **Objective:** Have more discs of your color than your opponent when the board is full or neither player can move. |
13 | | -- **Valid Move:** A move is made by placing a disc of your color on an empty square that "outflanks" one or more opponent discs. "Outflanking" means having a disc of your color at each end of a line (horizontal, vertical, or diagonal) of one or more contiguous opponent discs. |
14 | | -- **Flipping:** All outflanked opponent discs are flipped to your color. |
15 | | -- **Skipping Turns:** If a player has no valid moves, their turn is skipped. If neither player can move, the game ends. |
| 11 | +1. **Start the backend server:** `docker compose up` |
| 12 | +2. **Prepare the environment:** |
| 13 | + ```bash |
| 14 | + python3 -m venv venv |
| 15 | + source venv/bin/activate |
| 16 | + pip install -r requirements.txt |
| 17 | + ``` |
| 18 | +3. **Run the Main Agent (NNUE-DQN):** `python -m agents.ai_agent` |
| 19 | +4. *(Optional)* **Run the Classical Teacher:** `python -m agents.classical_agent -d h` |
16 | 20 |
|
17 | | -### State and Actions Example |
| 21 | +--- |
18 | 22 |
|
19 | | -The game state is communicated to agents as a JSON object: |
20 | | -```json |
21 | | -{ |
22 | | - "type": "state", |
23 | | - "board": [[0, 0, ...], ...], |
24 | | - "current_turn": 1, |
25 | | - "valid_actions": [[2, 3], [3, 2], [4, 5], [5, 4]] |
26 | | -} |
27 | | -``` |
28 | | -An action is a simple move command: |
29 | | -```json |
30 | | -{ |
31 | | - "action": "move", |
32 | | - "x": 2, |
33 | | - "y": 3 |
34 | | -} |
35 | | -``` |
| 23 | +## 2. Solution Architectures |
36 | 24 |
|
37 | | -## Setup |
| 25 | +Our group decided to build a highly optimized classical baseline to serve as both a rigorous benchmark and a "Teacher" for our modern Deep Learning approach. |
38 | 26 |
|
39 | | -The simulation can be launched using Docker Compose, while agents are typically executed locally. |
| 27 | +### 2.1. The Classical Baseline (Minimax + Numba) |
| 28 | +To establish a robust benchmark, we built a Minimax algorithm with Alpha-Beta pruning. |
| 29 | +* **Performance Optimization:** We flattened the 2D boards and compiled the core Othello logic into machine code using **Numba (`@njit`)**, achieving a ~50x speedup in evaluation. |
| 30 | +* **Heuristics & Move Ordering:** The agent evaluates positional weights and dynamic mobility, optimizing Alpha-Beta cutoffs by testing corner moves first. |
| 31 | +* **The "Predictable Teacher" Limitation:** While highly optimized, this Minimax implementation is purely deterministic. It will always play the exact same optimal sequence for any given board. As detailed in Section 3, this lack of stochasticity presented a significant challenge during the AI training phase. |
40 | 32 |
|
41 | | -1. **Start the Simulation**: |
42 | | - ```bash |
43 | | - docker compose up |
44 | | - ``` |
45 | | - This will start the backend server (port 8765) and the frontend viewer (port 8080). |
| 33 | +### 2.2. Residual Value Network (NNUE) with DQN |
| 34 | +Our final and best-performing agent utilizes a custom Neural Network architecture inspired by Efficiently Updatable Neural Networks (NNUE). Instead of predicting the action directly, the model acts as a **Value Network**, evaluating the strength of a given board state. |
46 | 35 |
|
47 | | -2. **Execute Agents**: |
48 | | - Create and activate a virtual environment, install the requirements, and run your agents: |
49 | | - ```bash |
50 | | - python3 -m venv venv |
51 | | - source venv/bin/activate |
52 | | - pip install -r requirements.txt |
53 | | - python agents/dummy_agent.py |
54 | | - ``` |
| 36 | +#### Feature Extraction (State Representation) |
| 37 | +Rather than feeding raw 8x8 grids, the board is transformed into a **132-dimensional feature vector**: |
| 38 | +* `[0:64]`: Agent's piece positions (1.0 if present, else 0.0). |
| 39 | +* `[64:128]`: Opponent's piece positions (1.0 if present, else 0.0). |
| 40 | +* `[128:132]`: Hand-crafted strategic heuristics normalized between -1 and 1: |
| 41 | + - Relative Corner Control. |
| 42 | + - X-Square Risk Penalty (avoiding corners' adjacent cells). |
| 43 | + - Relative Mobility (difference in available legal moves). |
| 44 | + - Center Control. |
| 45 | + |
| 46 | +#### Network Architecture (ResNet) |
| 47 | +The model was built using PyTorch and employs modern Deep Learning stabilization techniques: |
| 48 | +* **Input Layer:** A fully connected layer expanding the 132 features to 256 dimensions, followed by ReLU, **Layer Normalization**, and **Dropout (10%)** to prevent overfitting. |
| 49 | +* **Residual Block:** A hidden block with two 256-neuron linear layers using a **Skip Connection** (`x = res_block(x) + identity`). This residual architecture allows deeper feature correlation while mitigating the vanishing gradient problem. |
| 50 | +* **Output Head:** A bottleneck progression `(256 -> 64 -> 16 -> 1)` that condenses the spatial and strategic features into a single scalar value representing the state's Q-value. |
| 51 | + |
| 52 | +#### Deliberation Strategy |
| 53 | +During gameplay, the agent identifies all valid moves, simulates the board state resulting from each move, and evaluates them through the Residual Value Network. The move that yields the highest predicted state value is executed. |
55 | 54 |
|
56 | | -## Project Structure |
| 55 | +--- |
57 | 56 |
|
58 | | -- `backend/`: Python server using `websockets` that handles game logic and state broadcasting. |
59 | | -- `frontend/`: HTML5 Canvas-based visualization for monitoring the game. |
60 | | -- `agents/`: Framework and implementations for autonomous agents. |
61 | | - - `base_agent.py`: Abstract base class for all agents. |
62 | | - - `dummy_agent.py`: Simple agent that makes random moves. |
63 | | - - `manual_agent.py`: Agent for manual control via terminal input. |
64 | | -- `compose.yml`: Docker Compose configuration for the full stack. |
| 57 | +## 3. Engineering Challenges & Solutions |
65 | 58 |
|
66 | | -## Development |
| 59 | +Developing a generalized AI for Othello presented a major technical hurdle regarding how the agent generalized its knowledge. |
67 | 60 |
|
68 | | -To develop a new agent, inherit from `BaseOthelloAgent` and implement the `deliberate` method. |
| 61 | +### The "Bad Teacher" Problem (Deterministic Overfitting) |
| 62 | +Initially, we trained our neural network against our standard Minimax agent. The AI quickly achieved a 100% win rate during training but completely failed during real-world testing. |
69 | 63 |
|
70 | | -```python |
71 | | -from agents.base_agent import BaseOthelloAgent |
| 64 | +**The Cause:** Because the Minimax agent was purely deterministic, it acted as a predictable and inflexible teacher. The neural network did not learn the generalized rules of Othello; instead, it memorized a single, highly specific choreographed sequence of moves to exploit the Minimax's exact heuristic. As soon as a real match deviated by a single move, the AI's strategy collapsed. |
72 | 65 |
|
73 | | -class MyAgent(BaseOthelloAgent): |
74 | | - async def deliberate(self, board, valid_actions): |
75 | | - # Your strategy here |
76 | | - if valid_actions: |
77 | | - return valid_actions[0] |
78 | | - return None |
79 | | -``` |
| 66 | +**The Solution (Stochastic Openings):** We injected chaos into the curriculum by forcing the "Teacher" agent to play its first two moves completely at random. This TCEC-style (Top Chess Engine Championship) approach forced the Neural Network to start matches from thousands of unique, unpredictable board states, breaking the memorization loop and forcing true spatial generalization. We also implemented **Reward Shaping**, penalizing the agent for playing in hazardous X-Squares and rewarding it heavily for securing corners during training. |
80 | 67 |
|
81 | | -Refer to the [API Documentation](https://mariolpantunes.github.io/si2-othello/) for more details. |
| 68 | +--- |
82 | 69 |
|
83 | | -## Authors |
| 70 | +## 4. Final Performance Evaluation |
84 | 71 |
|
85 | | -* **Mário Antunes** - [mariolpantunes](https://github.com/mariolpantunes) |
| 72 | +To rigorously test our final NNUE-DQN agent, we ran a TCEC-style benchmark using 10 stochastic openings (playing each as both Black and White) to prevent deterministic sequence memorization. |
86 | 73 |
|
87 | | -## License |
| 74 | +| Opponent | Win Rate | Margin (Avg Pieces) | Conclusion | |
| 75 | +| :--- | :--- | :--- | :--- | |
| 76 | +| **Minimax Easy (Depth 2)** | **70.0%** | +7.2 | Agent consistently avoids shallow traps and dominates basic lookahead. | |
| 77 | +| **Minimax Normal (Depth 4)** | **50.0%** | +1.2 | Agent performs on par with a 4-step exhaustive search, proving the strategic depth of the NNUE features. | |
| 78 | +| **Minimax Hard (Depth 6 + Mob)** | **~10.0%** | Negative | Exhaustive deep search with mobility heuristics outperforms our model's immediate pattern recognition. | |
88 | 79 |
|
89 | | -This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details. |
| 80 | +**Final Conclusion:** |
| 81 | +The NNUE architecture proved to be highly effective and extremely fast at inference time. Achieving a 50% win rate against a robust Depth-4 Alpha-Beta search demonstrates that the network successfully learned deep spatial and mobility concepts (such as corner control and edge stability) without needing to explicitly traverse a complex decision tree. Furthermore, the development of the Numba-compiled classical engine was vital to provide a challenging training curriculum and properly benchmark the agent's limitations. An improvement that could be made was simply adding a discount factor, make the teacher have some randomness and not be greedy. |
0 commit comments