Optimizer research candidates

From The Hei Canon

Optimizer research candidates collects the optimizer alternatives discussed in the project beyond the implemented AdamW/Lion selection and per-objective isolation.

Project status: Mixed proposal/deferred status; none of the alternatives below is a registered choice in the audited Trainfer optimizer selector. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

The optimizer notes distinguish several levers:

  • Schedule-free optimization: remove dependence on a hand-designed finite training schedule for a continuing stream.
  • AdEMAMix8bit: add a slower momentum timescale as a candidate response to nonstationary training.
  • Muon/Riemannion: matrix-update geometry alternatives requiring care around parameter shape and model family.
  • Gradient surgery (PCGrad, CAGrad, GradVac, MGDA): compute and reconcile separate task gradients rather than merely rescale a summed loss.
  • EMA loss normalization/DB-MTL-style log transforms: normalize objective magnitudes before mixing them, without storing separate per-task gradients.

Implementation and controls

The prioritized optimizer note and production roadmap propose wrappers and controlled 500-event stream comparisons. The AdEMAMix spike sets a wall-time regression gate of 5%. The notes defer expensive gradient surgery because it needs separate backward passes and gradient storage. Current TrainEngine supports AdamW8bit/Lion8bit selection with plain AdamW fallbacks, plus separate plain AdamW instances.

Evidence and evaluation

These are research candidates and structural arguments in the cited notes. Their external benchmarks are not local efficacy evidence. No completed, comparable multi-arm result for all of these optimizers was found in the inspected reports.

Limitations and interpretation

Do not equate loss normalization with optimizer-history isolation or gradient-direction conflict resolution. A streaming daemon lacks a natural end-of-training checkpoint; methods with train/eval parameter averaging need an explicit serving policy. Model family, tensor shapes, memory and latency can rule out an otherwise plausible optimizer.

Sources

See also