Rehearsal and snapshot self-distillation
Rehearsal and snapshot self-distillation covers the project's proposed retention-focused training pulses and teacher-snapshot consolidation.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Rehearsal periodically retrains on examples from a skill that should be retained. Snapshot self-distillation instead generates targets or logits from a frozen previously useful model state and trains the live model toward them. Both differ from recency-weighted replay of raw user feedback.
Implementation and controls
Roadmap PR H refers to a rehearsal loop and its interaction with separate optimizer histories. Plan T4.2 proposes generation under a frozen best-adapter snapshot followed by KL-based self-distillation. The current frozen-reference loader loads the base model, not a best-adapter snapshot; that infrastructure alone is not the complete T4.2 loop.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
Rehearsing one skill may overfit its exemplars or interfere with new facts. A teacher snapshot can preserve obsolete mistakes. Compare heldout retention curves and adaptation after equal training tokens; existing idle feedback replay should not be relabelled as a completed snapshot-distillation system.
Sources
- cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - cont: docs/research/optimizer-sample-efficiency.md — checkout audited
87946914c7b9. - trainfer: trainfer/PLAN.md — checkout audited
1c6391f3773b. - trainfer: trainfer/state.py — checkout audited
1c6391f3773b.