Skip to content

Stateful step optimisers (SGD+momentum, Adam) for mini-batch training #339

Description

@gabrielfrasantos

Motivation

numerical/optimization only offers full-batch Optimizer::Minimize(initialGuess, objective), implemented by GradientDescent and BayesianOptimization. On-device neural-network training in neural-network-toobox-cpp (roadmap N14 → N16) needs an optimiser that takes one externally computed gradient per step and keeps its own state between steps.

Proposed scope

  • A StepOptimizer<T, N> interface: Step(θ, g) updates θ in place from an externally computed gradient, plus Reset().
  • SGD with optional momentum/Nesterov: N extra floats of state.
  • Adam with bias correction: 2N extra floats, with β₁, β₂, ε and the learning rate as parameters.
  • Float-only, no heap, fixed-size state in math::Vector<T, N>, OPTIMIZE_FOR_SPEED on Step.
  • Tests: one step against hand-computed values, convergence on a quadratic, Adam's bias correction at t = 1, and Reset() restoring the initial state.

References

  • B. Polyak, 1964.
  • I. Sutskever et al., ICML, 2013.
  • D. Kingma, J. Ba, "Adam," ICLR, 2015.

Downstream consumer: neural-network-toobox-cpp ROADMAP.md N14/N16 (spec roadmap/model/MiniBatchTraining/).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions