Trace infilling

From The Hei Canon

Trace infilling is the project's selective-supervision mechanism for preserving an accepted response prefix while training its continuation.

Project status: Prefix masking is implemented; the plan's richer critique-targeted multi-candidate infilling is a broader design. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Mark prompt tokens and an accepted span_prefix as ignored labels; compute SFT only after that prefix. This concentrates direct token supervision on the regenerated suffix. It is a loss-mask operation, not an attention mask that prevents the suffix from reading its prefix.

Implementation and controls

build_chat_inputs in objectives/_utils.py implements mask construction. SFT and weighted SFT pass the sample's span_prefix through. The design's T3.1 envisioned auxiliary sampling restricted to a critiqued reasoning span, but the audited CCPD sample schema does not wire a general span-infilling workflow. Entity-masked SFT is a concrete use of the existing prefix mask.

Evidence and evaluation

The source includes test_span_prefix.py for prompt/prefix masking and suffix supervision. This tests geometry, not a claim that untargeted model behavior never changes. The entity-masked experiment provides a small recipe-level ablation.

Limitations and interpretation

Tokenization boundaries and chat closing markers affect which tokens are actually supervised. An entirely masked response yields no useful target loss. Masking token losses does not freeze shared weights or remove effects on other prompts. The planned “almost no side effects” statement is not supported by mask tests alone.

Sources

See also