Rehearsal and snapshot self-distillation

From The Hei Canon

Rehearsal and snapshot self-distillation covers the project's proposed retention-focused training pulses and teacher-snapshot consolidation.

Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Rehearsal periodically retrains on examples from a skill that should be retained. Snapshot self-distillation instead generates targets or logits from a frozen previously useful model state and trains the live model toward them. Both differ from recency-weighted replay of raw user feedback.

Implementation and controls

Roadmap PR H refers to a rehearsal loop and its interaction with separate optimizer histories. Plan T4.2 proposes generation under a frozen best-adapter snapshot followed by KL-based self-distillation. The current frozen-reference loader loads the base model, not a best-adapter snapshot; that infrastructure alone is not the complete T4.2 loop.

Evidence and evaluation

The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.

Limitations and interpretation

Rehearsing one skill may overfit its exemplars or interfere with new facts. A teacher snapshot can preserve obsolete mistakes. Compare heldout retention curves and adaptation after equal training tokens; existing idle feedback replay should not be relabelled as a completed snapshot-distillation system.

Sources

See also