Lion8bit optimization
Lion8bit optimization is an optional optimizer choice evaluated as an alternative to shared AdamW moment scaling in the project's continual-learning design.
Project status: Selection implemented as optimizer="lion8bit"; not the default. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
The project delegates the actual Lion update to bitsandbytes rather than defining a custom loss or optimizer here. The research argument is to avoid Adam's shared second-moment denominator when objective scales differ. This changes parameter-update dynamics while leaving data and objective construction intact.
Implementation and controls
TrainEngine._optimizer selects bnb.optim.Lion8bit in shared mode when requested. Import/initialization failure falls back to torch AdamW and logs a warning. When per-objective isolation is enabled, the code chooses plain AdamW instances instead, so requesting Lion alongside that flag does not establish a Lion run.
Evidence and evaluation
The optimizer research note contains structural motivation and an A/B proposal. The presence of the historical branch name feat/lion8bit-optimizer also preserved V-KTO history, but reachability through that branch does not imply V-KTO was trained with Lion. Actual configuration and logs must establish the optimizer for each run.
Limitations and interpretation
Avoid describing unavailable bitsandbytes fallback as a successful Lion experiment. Learning-rate scales are not necessarily transferable from AdamW. No performance or forgetting advantage follows solely from eliminating a second-moment accumulator; compare convergence, retention, latency, and memory on the same stream.
Sources
- trainfer: trainfer/engine/train.py — checkout audited
1c6391f3773b. - trainfer: trainfer/config.py — checkout audited
1c6391f3773b. - cont: docs/research/optimizer-sample-efficiency.md — checkout audited
87946914c7b9.