Build a Trainable CNN from Scratch
Assemble a LeNet-style convolutional network with im2col convolutions, gradients, Adam, and a complete training loop.
Begin with the problem, not the library
Before Build a Trainable CNN from Scratch is a collection of classes and functions, it is an answer to a constraint. Assemble a LeNet-style convolutional network with im2col convolutions, gradients, Adam, and a complete training loop. The useful question is not “which API should I call?” but “what information is available, what decision must be made, and what evidence proves the decision is good?”
A first-principles implementation makes hidden assumptions visible. It forces us to specify the input, the transformation, the objective, and the failure conditions. That discipline is valuable even when a production system later uses a mature library.
Reduce the system to four questions
Representation
How is the raw problem expressed as numbers, states, tokens, tensors, or events?
Objective
What quantity tells the system that one answer is better than another?
Update
How does evidence change parameters, state, policy, or decisions?
Evaluation
Which controlled test separates real improvement from noise or leakage?
Build a Trainable CNN from Scratch becomes understandable when each implementation step answers exactly one of these questions. The walkthrough keeps those boundaries explicit so a bug can be localized instead of disappearing inside an end-to-end pipeline.
The ideas you must genuinely understand
im2col
A convolution shares a local kernel across spatial positions. im2col unfolds receptive fields into matrix columns so the forward pass becomes GEMM; the backward pass must scatter overlapping column gradients back to the image.
In Build a Trainable CNN from Scratch, implement this idea first on a tiny hand-computable example. Write down every shape, legal range, and invariant; compare the code with the manual result; then profile and scale only after the reference agrees.
Verification rule: test the normal case, a boundary case, an invalid case, and an invariant that must remain true after the operation.
Pooling
Pooling summarizes a local neighborhood to reduce spatial resolution. Max pooling routes gradient only to the recorded argmax, making index caching and tie behavior important for a correct backward pass.
In Build a Trainable CNN from Scratch, implement this idea first on a tiny hand-computable example. Write down every shape, legal range, and invariant; compare the code with the manual result; then profile and scale only after the reference agrees.
Verification rule: test the normal case, a boundary case, an invalid case, and an invariant that must remain true after the operation.
Adam
Adam tracks exponential moving averages of gradients and squared gradients, applies bias correction early in training, and scales each parameter update by its estimated second moment. Epsilon placement and weight-decay semantics matter.
In Build a Trainable CNN from Scratch, implement this idea first on a tiny hand-computable example. Write down every shape, legal range, and invariant; compare the code with the manual result; then profile and scale only after the reference agrees.
Verification rule: test the normal case, a boundary case, an invalid case, and an invariant that must remain true after the operation.
From first principles to production evidence
The following chapters deliberately slow the build down. They connect every major milestone to its contract, derivation, implementation choices, tests, failure modes, systems cost, and production responsibilities.
Verified as part of a 10,000+ word project articleFormulate the problem before choosing the machinery
Build a Trainable CNN from Scratch begins with a decision problem, not a framework. Assemble a LeNet-style convolutional network with im2col convolutions, gradients, Adam, and a complete training loop. Restate that sentence as an observable input, a desired output, and a criterion for preferring one output over another. Identify who or what supplies supervision, whether feedback is immediate or delayed, and whether examples can be considered independent. These choices determine what can be learned and what remains an assumption. The implementation is honest only when those assumptions are visible near the data contract rather than buried in training code.
The raw material becomes a feature representation. Representation decides which distinctions the system can express and which distinctions disappear. List categorical domains, numerical units, missing-value semantics, sequence or spatial axes, masks, player or client perspective, and precision. Then consider invariances: should translation, permutation, rescaling, token position, client identity, or board symmetry change the answer? An architecture that ignores the required invariance wastes data; one that imposes the wrong invariance makes the target impossible to represent.
Finally define the baseline and the abstention point. A baseline can be a constant predictor, random policy, linear rule, naive kernel, synchronous algorithm, or human heuristic. It anchors complexity in evidence. The abstention point describes inputs for which the system lacks support and should decline, defer, or fall back. Together they prevent Build a Trainable CNN from Scratch from being judged only by an impressive end-to-end demonstration while basic correctness, calibration, robustness, or operational usefulness remains unknown.
Connect the objective to the behavior you actually want
An objective compresses preferences into a scalar, but no scalar captures every product or scientific goal. For Build a Trainable CNN from Scratch, distinguish the training objective from the evaluation metric and the deployment utility. The training objective must provide a usable signal to parameters or state; evaluation must estimate generalization under a controlled protocol; deployment utility includes latency, cost, safety, and the consequence of errors. When these three disagree, optimization can succeed while the system becomes less useful.
Study each term dimensionally and statistically. Ask what happens if one term is multiplied by ten, one class becomes rare, a sequence becomes longer, a client contributes more samples, or rewards are shifted. Determine whether averages are per token, example, client, action, spatial position, or batch. Regularization is not decorative: it encodes a preference over solutions and changes units unless normalized consistently. A correct derivation names the population quantity of interest, its finite-sample estimator, and the approximation introduced by minibatches, replay, sampling, or surrogate losses.
Identifiability is the deeper constraint. Data may not contain enough information to separate competing explanations. im2col, Pooling, Adam can improve computation or inductive bias, but they cannot manufacture missing evidence. State causal assumptions, observability limits, support conditions, and equivalence classes of solutions. Use sensitivity analysis and targeted interventions where possible. When identification is impossible, report uncertainty or a set of plausible answers rather than converting an arbitrary modeling choice into unwarranted confidence.
Make mathematical equivalence survive finite precision
Paper algebra assumes exact real numbers; the implementation uses finite precision, bounded memory, and discrete execution order. In Build a Trainable CNN from Scratch, audit exponentials, logarithms, divisions, reductions, norms, probabilities, recursive values, and accumulated updates. Rewrite unstable expressions with max subtraction, log-sum-exp, compensated accumulation, safe denominators, or higher-precision reductions. Track where a mathematically harmless reordering changes rounding and where mixed precision needs scaling or master copies.
Shapes are part of the proof. Annotate each intermediate with semantic axes rather than only dimensions: batch, token, head, channel, client, action, expert, feature, row, column, or sample. Broadcasting should be intentional and verified with asymmetric dimensions so an accidental match cannot hide. Record contiguous layout and stride assumptions when performance code depends on them. For every reshape or transpose, write both the precondition and the inverse operation needed during backward, decoding, aggregation, or reconstruction.
Build a numerical ladder: scalar example, tiny vector or matrix example, batched reference, optimized path, then realistic workload. At each rung compare values and invariants before increasing scale. This catches defects while they are still interpretable. The acceptance test should specify absolute and relative error, exceptional values, deterministic modes, and the hardware or library versions used. Numerical stability is not a final cleanup task; it is part of the algorithm’s definition.
Design evidence that can falsify the implementation
Evaluation is an experiment. For Build a Trainable CNN from Scratch, specify the unit of analysis, split strategy, temporal boundary, randomization, baseline, metric, and uncertainty before viewing final results. Prevent duplicates, transformed copies, future information, opponent leakage, and shared-client information from crossing the boundary. A single aggregate score can hide subgroup collapse, unstable seeds, poor calibration, tail latency, or rare catastrophic behavior, so pair it with distributions and stratified slices.
Ablations connect outcomes to mechanisms. Remove or replace im2col, Pooling, Adam one at a time while controlling data, compute, and evaluation. Compare equal wall-clock or equal resource budgets when efficiency is part of the claim. Repeat stochastic runs and report variation rather than selecting the best seed. Inspect learning curves and intermediate metrics because two systems with the same final score may differ radically in sample efficiency, stability, or cost.
The test suite and the benchmark answer different questions. Unit and property tests prove local contracts; integration tests prove components agree; benchmarks estimate behavior at scale; task evaluation estimates usefulness. Preserve all four. A benchmark that bypasses validation or uses a different code path from production is weak evidence. The strongest release gate reruns the exact packaged implementation with recorded configuration and produces an artifact that another person can inspect.
Turn the learning artifact into an operable system
Production structure separates pure computation from orchestration, configuration, persistence, and interfaces. Package the core of Build a Trainable CNN from Scratch behind typed contracts. Keep data loading, model or state construction, training, evaluation, serialization, and serving independently invocable. Configuration should be validated, versioned, and printable. Random seeds, data identifiers, source commit, dependency lock, hardware, and metric definitions belong in the run record so an apparent regression can be reproduced instead of guessed at.
Capacity planning follows the critical path. Measure sample complexity, arithmetic work, memory, and validation effort across representative input sizes and concurrency. Report warm-up separately, distinguish throughput from latency, and include tail percentiles. Define memory ownership and lifetime so caches, activations, buffers, replay, or optimizer state cannot grow without a bound. Backpressure and admission control are preferable to unpredictable collapse. Where hardware-specific acceleration exists, preserve a portable reference path for correctness and degraded operation.
Observability must explain decisions and failures without exposing sensitive content. Log stable identifiers, shapes, versions, summary statistics, timings, and error categories. Monitor input drift, output distribution, task quality, saturation, retries, and fallback rate. Establish rollback and shadow-evaluation procedures before the first risky change. A production-grade implementation is not merely more abstract than a notebook; it makes dependencies, state, failure, and evidence explicit enough for another engineer to operate safely.
Read claims as reproducible hypotheses
The research surrounding Build a Trainable CNN from Scratch improves representations, objectives, algorithms, systems, or evaluation protocols. Classify each paper by which lever it changes. Then identify the comparison budget: data, parameters, tokens, environment steps, hardware, communication, wall-clock time, and tuning effort. A claimed improvement may disappear when budgets are normalized or when the baseline receives equal tuning. Read methods and appendices for details that determine reproducibility, not only the abstract and headline table.
Reproduction begins with the smallest claim. Recreate one table row or ablation before attempting the entire system. Preserve the authors’ preprocessing and metric definitions, then deliberately vary one assumption. Document deviations, failed attempts, and environment details. When a result does not reproduce, distinguish an implementation defect from missing procedural knowledge, stochastic uncertainty, and genuine sensitivity. Negative evidence is useful when it narrows the conditions under which the method works.
Extension should start from a mechanism and a falsifiable prediction. The skills developed here—Convolutions, Backpropagation, Training loops—suggest multiple directions, but change one major factor at a time. Predict which metric and intermediate signal should move if the explanation is correct. Use confidence intervals and preregistered stopping rules for expensive experiments where possible. Publish code, configuration, data provenance, and failure cases so the work contributes more than another isolated score.
Maintain a chain of evidence from equation to outcome
A proof ledger for Build a Trainable CNN from Scratch links each important claim to the smallest evidence that could disprove it. For a mathematical claim, keep a hand-worked example and a high-precision reference. For a software contract, keep unit and property tests. For an optimization claim, keep profiler traces and equal-budget baselines. For a learning claim, keep per-seed results, confidence intervals, and ablations. For a production claim, keep load tests, failure injection, monitoring queries, and rollback evidence. This structure prevents one successful end-to-end run from being treated as proof of every layer beneath it.
Record evidence beside the versioned artifact it evaluates. A metric without its dataset revision, configuration, dependency lock, hardware, and commit cannot reliably settle a regression. Likewise, a screenshot or generated sample is qualitative evidence, not a distribution. Name the claim, evidence type, acceptance threshold, owner, and date. When the implementation changes, rerun the smallest affected evidence first and then the downstream integration gates. The ledger becomes a map of confidence: it shows what is known, what is assumed, what has become stale, and where another experiment is required.
Use the ledger during review. Ask whether each test would fail for a realistic defect, whether each benchmark measures the packaged code path, whether every aggregate retains inspectable raw values, and whether uncertainty is reported at the correct independent unit. Include counterexamples and failed experiments because they define the boundary of the method. Over time this habit turns Convolutions, Backpropagation, Training loops from isolated implementation skills into a reproducible engineering practice that survives new data, new hardware, new collaborators, and changing product constraints.
Argmax Rows — from contract to production evidence
Argmax Rows is the pipeline boundary at milestone 1 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between softmax, loss, and metrics primitives and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Argmax Rows as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Argmax Rows depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Argmax Rows needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Argmax Rows can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Argmax Rows changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Softmax, Loss, and Metrics Primitives. Build the numerically stable softmax, cross-entropy loss, and accuracy helpers used throughout the network.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Exp Shifted — from contract to production evidence
Exp Shifted is the pipeline boundary at milestone 4 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between softmax, loss, and metrics primitives and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Exp Shifted as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Exp Shifted depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Exp Shifted needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Exp Shifted can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Exp Shifted changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Softmax, Loss, and Metrics Primitives. Build the numerically stable softmax, cross-entropy loss, and accuracy helpers used throughout the network.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Gather True Class Probs — from contract to production evidence
Gather True Class Probs is the pipeline boundary at milestone 7 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between softmax, loss, and metrics primitives and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Gather True Class Probs as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Gather True Class Probs depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Gather True Class Probs needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Gather True Class Probs can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Gather True Class Probs changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Softmax, Loss, and Metrics Primitives. Build the numerically stable softmax, cross-entropy loss, and accuracy helpers used throughout the network.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
He Std — from contract to production evidence
He Std is the pipeline boundary at milestone 10 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between initialization and convolution plumbing and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat He Std as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If He Std depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for He Std needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how He Std can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing He Std changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Initialization and Convolution Plumbing. Implement He initialization, zero biases, padding, output-shape math, and the im2col / col2im transforms that power efficient convolutions.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Output Spatial Size — from contract to production evidence
Output Spatial Size is the pipeline boundary at milestone 14 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between initialization and convolution plumbing and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Output Spatial Size as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Output Spatial Size depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Output Spatial Size needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Output Spatial Size can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Output Spatial Size changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Initialization and Convolution Plumbing. Implement He initialization, zero biases, padding, output-shape math, and the im2col / col2im transforms that power efficient convolutions.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Conv2d Forward — from contract to production evidence
Conv2d Forward is the transformation at milestone 17 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Conv2d Forward as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Conv2d Forward depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Conv2d Forward needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Conv2d Forward can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Conv2d Forward changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Conv2d Grad Bias — from contract to production evidence
Conv2d Grad Bias is the learning update at milestone 20 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Conv2d Grad Bias as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Conv2d Grad Bias depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Conv2d Grad Bias needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Conv2d Grad Bias can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Conv2d Grad Bias changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Scatter Grad Window — from contract to production evidence
Scatter Grad Window is the learning update at milestone 23 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Scatter Grad Window as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Scatter Grad Window depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Scatter Grad Window needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Scatter Grad Window can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Scatter Grad Window changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Flatten Forward — from contract to production evidence
Flatten Forward is the transformation at milestone 27 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Flatten Forward as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Flatten Forward depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Flatten Forward needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Flatten Forward can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Flatten Forward changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Linear Grad Input — from contract to production evidence
Linear Grad Input is the learning update at milestone 30 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Linear Grad Input as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Linear Grad Input depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Linear Grad Input needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Linear Grad Input can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Linear Grad Input changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Linear Backward — from contract to production evidence
Linear Backward is the learning update at milestone 33 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between layer forward and backward passes and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Linear Backward as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Linear Backward depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Linear Backward needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Linear Backward can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Linear Backward changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Layer Forward and Backward Passes. Code the forward and backward routines for convolution, max pooling, ReLU, flatten, and linear layers.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Adam Update M — from contract to production evidence
Adam Update M is the learning update at milestone 37 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between fused loss and optimizers and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Adam Update M as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Adam Update M depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Adam Update M needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Adam Update M can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Adam Update M changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Fused Loss and Optimizers. Fuse softmax with cross-entropy for stable training and implement both SGD and Adam parameter updates.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Adam Param Step — from contract to production evidence
Adam Param Step is the learning update at milestone 40 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between fused loss and optimizers and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Adam Param Step as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Adam Param Step depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Adam Param Step needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Adam Param Step can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Adam Param Step changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Fused Loss and Optimizers. Fuse softmax with cross-entropy for stable training and implement both SGD and Adam parameter updates.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Init Linear Layer — from contract to production evidence
Init Linear Layer is the construction at milestone 43 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between assembling lenet and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Init Linear Layer as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Init Linear Layer depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Init Linear Layer needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Init Linear Layer can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Init Linear Layer changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Assembling LeNet. Compose the layer primitives into convolutional and classifier blocks, wire up the full LeNet forward/backward pass, and add a predict helper.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Forward Classifier Block — from contract to production evidence
Forward Classifier Block is the transformation at milestone 46 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between assembling lenet and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Forward Classifier Block as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Forward Classifier Block depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Forward Classifier Block needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Forward Classifier Block can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Forward Classifier Block changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Assembling LeNet. Compose the layer primitives into convolutional and classifier blocks, wire up the full LeNet forward/backward pass, and add a predict helper.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Lenet Backward — from contract to production evidence
Lenet Backward is the learning update at milestone 50 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between assembling lenet and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Lenet Backward as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Lenet Backward depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Lenet Backward needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Lenet Backward can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Lenet Backward changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Assembling LeNet. Compose the layer primitives into convolutional and classifier blocks, wire up the full LeNet forward/backward pass, and add a predict helper.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Shuffle Indices — from contract to production evidence
Shuffle Indices is the pipeline boundary at milestone 53 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between synthetic data pipeline and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Shuffle Indices as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Shuffle Indices depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Shuffle Indices needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Shuffle Indices can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Shuffle Indices changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Synthetic Data Pipeline. Generate a small synthetic image dataset and build shuffling, train/test splitting, and minibatch iteration utilities.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Train Step — from contract to production evidence
Train Step is the learning update at milestone 56 of Build a Trainable CNN from Scratch. Its purpose is not merely to make the next function run. It establishes a contract between training loop and evaluation and every downstream stage. Begin by naming the accepted inputs, their axes, units, legal ranges, ownership rules, and whether mutation is permitted. Then name the output with the same precision. In this project the surrounding ideas—im2col, Pooling, Adam—only compose correctly when this boundary preserves those invariants. A useful implementation note records one representative shape, one smallest valid example, one boundary example, and one invalid example before any optimization is attempted.
From first principles, treat Train Step as a mapping from available information to a new feature representation. Ask which information is genuinely known at this point and which information would leak from the future, evaluation set, opposing player, held-out client, or later pipeline stage. Write the transformation symbolically before translating it into array operations. Every reduction must state its axis; every probability must state its normalization set; every random choice must state its distribution and seed; every learned quantity must state the objective that changes it. This discipline turns an appealing formula into an executable specification that can be challenged with small counterexamples.
The reference implementation should favor clarity over cleverness. Separate validation, the mathematical core, and state updates so each can be tested independently. Use explicit intermediate names that correspond to the derivation rather than compressing the work into one expression. Confirm dtype promotion, broadcasting, device placement, and empty-input behavior. If Train Step depends on randomness, pass a generator instead of reading hidden global state. If it owns mutable state, return or document the updated state explicitly. The optimized implementation may later fuse operations or reuse buffers, but it must remain numerically comparable with this small version on deterministic fixtures.
Verification for Train Step needs more than a happy-path assertion. Prove a hand-computable normal case, a boundary case, an invalid case, and at least one invariant. Compare against hand calculations, unit tests, controlled baselines, and held-out metrics. Add metamorphic tests when an exact answer is awkward: permutation, scaling, symmetry, conservation, monotonicity, or equivalence under a harmless representation change. Run the test repeatedly under fixed seeds to distinguish deterministic defects from statistical variation. When floating-point arithmetic is involved, justify tolerances from expected rounding error instead of choosing a loose threshold simply because the test passes.
Failure analysis asks how Train Step can look plausible while being wrong. Inspect leakage, numerical instability, overfitting, shape errors, and misleading aggregate metrics. Trace one example through every intermediate value and preserve enough logging to reproduce it. Distinguish a contract violation from an optimization failure and from an evaluation-design failure; each requires a different repair. A numerical answer within range is not automatically meaningful, and a rising training metric is not proof that the intended signal is being learned. The strongest debugging move is usually to shrink the input until the complete computation fits on paper, then compare the paper trace with the program line by line.
Productionizing Train Step changes the question from “does it work once?” to “does it remain trustworthy under load and change?” Measure sample complexity, arithmetic work, memory, and validation effort. Define observability for inputs, outputs, latency, failures, drift, and resource saturation. Decide what happens on malformed data, cancellation, partial worker failure, unavailable accelerators, or a distribution outside the training envelope. Version configuration and schemas with the code, preserve reproducible seeds where appropriate, and expose a safe fallback. Optimization is accepted only when the reference tests, numerical comparisons, and task-level metrics remain within an explicitly documented budget.
- Part: Training Loop and Evaluation. Tie everything together with a training step, epoch loop, full training driver, and a held-out evaluation routine.
- Normal case: choose the smallest input that exercises the intended transformation.
- Boundary case: use an empty, singleton, saturated, masked, terminal, or maximum-size input as appropriate.
- Invariant: verify shape, range, conservation, normalization, symmetry, immutability, or monotonicity.
- Production evidence: record correctness, latency, memory or cost, and the exact configuration.
Where this pattern becomes useful
Convolutions
Use this capability when the product must make repeatable decisions under the same structural constraints studied in the project. Begin with an offline baseline, define a business-facing metric, and add monitoring before automation.
Use case 1Backpropagation
Use this capability when the product must make repeatable decisions under the same structural constraints studied in the project. Begin with an offline baseline, define a business-facing metric, and add monitoring before automation.
Use case 2Training loops
Use this capability when the product must make repeatable decisions under the same structural constraints studied in the project. Begin with an offline baseline, define a business-facing metric, and add monitoring before automation.
Use case 3How the field keeps improving
LeNet established end-to-end convolution, subsampling, and backpropagation; AlexNet showed that depth, ReLU, data augmentation, dropout, and GPU compute could scale CNNs; residual learning made substantially deeper networks trainable. A NumPy implementation should retain explicit convolution gradients while testing initialization, normalization, optimizer behavior, and vectorized im2col memory costs.
Improvements usually change one of four levers: representation, learning signal, computation path, or evaluation protocol. Read each source with its assumptions and comparison budget in view.
Gradient-Based Learning Applied to Document Recognition
Presented LeNet-style convolutional networks trained end-to-end with backpropagation for document recognition.
ImageNet Classification with Deep Convolutional Neural Networks
Demonstrated large-scale deep CNN training with GPUs, ReLU, augmentation, and dropout, decisively improving ImageNet classification.
Deep Residual Learning for Image Recognition
Introduced residual connections that make very deep convolutional networks easier to optimize.
Treat paper claims as hypotheses: reproduce the baseline, inspect ablations, normalize compute budgets, and verify whether the evaluation matches your intended use.
Your next-study roadmap
- Re-derive
Explain each core equation without looking at the code.
- Rebuild
Implement the smallest version again from an empty file.
- Stress test
Create adversarial, boundary, numerical, and distribution-shift tests.
- Read critically
Choose one foundational paper and two recent follow-ups; reproduce one reported comparison.
- Extend
Change one assumption, record the hypothesis, and run a controlled experiment.
- Publish
Document architecture, tradeoffs, failures, metrics, cost, and reproducible commands.