Lion8bit optimization

From The Hei Canon

Lion8bit optimization is an optional optimizer choice evaluated as an alternative to shared AdamW moment scaling in the project's continual-learning design.

Project status: Selection implemented as optimizer="lion8bit"; not the default. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

The project delegates the actual Lion update to bitsandbytes rather than defining a custom loss or optimizer here. The research argument is to avoid Adam's shared second-moment denominator when objective scales differ. This changes parameter-update dynamics while leaving data and objective construction intact.

Implementation and controls

TrainEngine._optimizer selects bnb.optim.Lion8bit in shared mode when requested. Import/initialization failure falls back to torch AdamW and logs a warning. When per-objective isolation is enabled, the code chooses plain AdamW instances instead, so requesting Lion alongside that flag does not establish a Lion run.

Evidence and evaluation

The optimizer research note contains structural motivation and an A/B proposal. The presence of the historical branch name feat/lion8bit-optimizer also preserved V-KTO history, but reachability through that branch does not imply V-KTO was trained with Lion. Actual configuration and logs must establish the optimizer for each run.

Limitations and interpretation

Avoid describing unavailable bitsandbytes fallback as a successful Lion experiment. Learning-rate scales are not necessarily transferable from AdamW. No performance or forgetting advantage follows solely from eliminating a second-moment accumulator; compare convergence, retention, latency, and memory on the same stream.

Sources

See also