LVL 01SK
Project overview
SYSTEM ARCHITECTURE

Tic-Tac-Toe: Minimax to DQN

Follow the data, decisions, feedback, and validation boundaries before writing the full system.

Tic-Tac-Toe: Minimax to DQN first-principles architecture infographic

How to read this diagram

Read left to right for the forward path: raw information becomes a representation, passes through the project’s main computational ideas, and produces an output that can be measured. Then follow the feedback path back toward the trainable or decision-making components.

01

DQN

DQN approximates action values with a neural network, trains from replay, and uses a slower target network to reduce moving-target instability. Legal-action masking, exploration, and overestimation bias are central details.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

02

Bootstrapping

Bagging trains estimators on bootstrap samples drawn with replacement. Each learner sees a slightly different empirical distribution. Averaging their predictions keeps shared signal while cancelling part of their uncorrelated variance.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

03

Exploration

LoRA freezes a base weight matrix and learns a low-rank update BA. The rank limits trainable capacity and memory cost; placement, scaling, initialization, and target modules determine whether the adapter can express the required behavior.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

Architecture review checklist

  • Every arrow has a documented shape, dtype, unit, or schema.
  • Training and evaluation paths cannot leak information into each other.
  • Randomness is seeded and captured in experiment metadata.
  • Expensive stages expose timing, memory, throughput, and error metrics.
  • Each feedback loop has a stop condition and a rollback strategy.
  • Small reference implementations exist for numerical comparisons.