Context distillation (Trainfer)
Context distillation (Trainfer) is the ccd objective intended to internalize supplied context so later queries can omit it. It is distinct from critique-conditional CCPD.
Project status: Registered as ccd; “fact-preserving” is its design goal, not a proven guarantee. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Run a frozen adapter-disabled teacher on context plus prompt, and the trainable student on the prompt alone. Align a trailing teacher span to the student span and minimize KL(teacher distribution || student distribution). Optionally match hidden states and add response SFT. For samples carrying fact probes, generate a student answer, check listed facts, and multiply that sample's KL by 1.5 minus its fact-retention score.
Implementation and controls
objectives/ccd.py::ccd_loss takes context, prompt, optional response, and optional probe_kind="fact"/facts. It requires disable_adapter(). Defaults: KL weight 1, hidden matching enabled, hidden weight 0.5, last four hidden layers, SFT weight 1.
Teacher/student inputs use chat templates when available. KL supervision covers aligned prompt/probe positions, not a teacher-generated answer distribution over an entire response. A provided response creates a separate student SFT term.
Evidence and evaluation
The current source establishes the implementation and its telemetry (ccd_kl, hidden loss, fact retention). The inspected research journals do not supply a controlled CCD retention curve comparable to the memorize or V-KTO campaigns. Context-distillation survey material supplies motivation rather than local performance proof.
Limitations and interpretation
The alignment assumes the final student-length span of teacher input corresponds to the student sequence. Context concatenation and chat-template wrappers make this an assumption worth testing per tokenizer. Matching prompt states does not guarantee accurate context-free answers. Fact-probe string checks are limited evidence; adapter-disabled teachers can still include merged residual effects. Additional teacher forwards and hidden states impose real memory/compute costs.
Sources
- trainfer: trainfer/objectives/ccd.py — checkout audited
1c6391f3773b. - trainfer: trainfer/objectives/verifiers/fact_verifier.py — checkout audited
1c6391f3773b. - cont: docs/research/surveys/context-distillation-2025.md — checkout audited
87946914c7b9.