OceanGo
A deep reinforcement learning Go engine recreating the foundational AlphaGo and AlphaZero architectures, engineered with centralized GPU inference queues to train and run efficiently on consumer hardware.
Key Architectural Features
19-Block ResNet Backbone
Deep convolutional residual network with 256 feature channels evaluating board positions simultaneously across a Policy Head (move probabilities) and a Value Head (win/loss evaluation).
Centralized GPU Inference Queue
Client/server multiprocessing pipeline where multiple CPU MCTS workers submit positions to a single dedicated GPU batch thread. Bypasses Python GIL and completely prevents Out-Of-Memory (OOM) crashes.
Batched MCTS with Virtual Loss
Monte Carlo Tree Search with PUCT (Predictor Upper Confidence Bounds) running 800 simulations per move. Applies virtual loss across parallel search paths to ensure diverse tree exploration.
17-Plane Board State History
Feeds the neural network the previous 8 board states along with color perspective channels, giving the engine an explicit spatial representation of situational tactics, momentum, and the Ko rule.
Self-Play Data Augmentation
8-fold symmetrical rotation and reflection matrix transformations on self-play boards, Dirichlet noise exploration injection, and temperature decay from exploration to exploitation.
Interactive Pygame UI
Real-time graphical board interface with mouse-driven move placement, turn pass support (via key shortcut), last-move markers, and instant AI tree search response display.
System Architecture & Specification
ResNet-19 Dual-Head Network
- Residual Blocks: 19 convolutional residual blocks with 256 feature channels and Batch Normalization.
- Input Representation: 17 planes representing the last 8 turns per player + turn indicator.
- Dual Output Heads: Policy head (move probability distribution) + Value head (game win/loss evaluation).
Batched PUCT Tree Search
- Simulations: 800 MCTS expansions per move using Predictor Upper Confidence Bounds (c_puct = 1.0).
- Exploration Noise: Root Dirichlet noise injection (α = 0.03, ε = 0.25) to explore novel tactical lines.
- Virtual Loss: Parallel thread path locking preventing multiple CPU workers from exploring identical nodes.
Training & Data Pipeline
- Optimizer: SGD with momentum (0.9), learning rate 0.01, and L2 weight decay penalty.
- Experience Buffer: 500,000 moves circular replay buffer preventing overfitting to recent games.
- Data Augmentation: 8-fold dihedral group symmetry transformations (rotations & reflections).
Source Code, Training Guides & Checkpoints
The full Python implementation, PyTorch model checkpoints, training scripts, and interactive Pygame game loop are available in the open-source repository on GitHub. Refer to the repository documentation for dependency setup and training parameters.
Frequently Asked Questions & Reinforcement Learning Architecture
What is OceanGo and how does it implement AlphaZero?
OceanGo is a deep reinforcement learning Go engine built in Python using PyTorch. It recreates AlphaZero using a 19-block Residual Network with dual Policy (move probability) and Value (win probability) output heads, 17-plane board history encoding, and PUCT Monte Carlo Tree Search.
How does the centralized GPU inference queue work?
Multiple CPU worker processes perform parallel MCTS tree traversals and submit board states to a single centralized GPU batch evaluation thread. This bypasses the Python Global Interpreter Lock (GIL) and prevents Out-Of-Memory (OOM) errors during self-play.
What search parameters are configured in OceanGo?
OceanGo runs 800 MCTS simulations per move using Predictor Upper Confidence Bounds (c_puct = 1.0), root Dirichlet noise exploration (alpha = 0.03, epsilon = 0.25), and virtual loss path locking to ensure multi-threaded exploration diversity.
Is there an interactive interface to play against the AI?
Yes. OceanGo includes an interactive graphical Pygame board interface supporting Human vs. AI matches with mouse click move placement, pass controls, last-move highlights, and real-time AI win rate evaluations.