LVL 01SK
Project overview
SYSTEM ARCHITECTURE

Q-Learning on FrozenLake

Follow the data, decisions, feedback, and validation boundaries before writing the full system.

Q-Learning on FrozenLake first-principles architecture infographic

How to read this diagram

Read left to right for the forward path: raw information becomes a representation, passes through the project’s main computational ideas, and produces an output that can be measured. Then follow the feedback path back toward the trainable or decision-making components.

01

Bellman update

The Bellman relationship decomposes long-term value into immediate reward plus discounted value of what follows. Bootstrapping replaces the unknown future return with the learner’s current estimate, which improves data efficiency but can amplify estimation error.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

02

Epsilon greedy

Epsilon greedy defines one of the project’s main information transformations. Understand its input representation, objective, numerical invariants, computational cost, and failure modes before relying on a library implementation.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

03

Q-table

The Bellman relationship decomposes long-term value into immediate reward plus discounted value of what follows. Bootstrapping replaces the unknown future return with the learner’s current estimate, which improves data efficiency but can amplify estimation error.

Boundary check: document its accepted input, output shape, mutable state, failure modes, and the metric that proves this stage is correct before connecting it downstream.

Architecture review checklist

  • Every arrow has a documented shape, dtype, unit, or schema.
  • Training and evaluation paths cannot leak information into each other.
  • Randomness is seeded and captured in experiment metadata.
  • Expensive stages expose timing, memory, throughput, and error metrics.
  • Each feedback loop has a stop condition and a rollback strategy.
  • Small reference implementations exist for numerical comparisons.