Stacked training recipes

From The Hei Canon

Stacked training recipes compose existing cont training mechanisms sequentially inside one experiment.

Project status: Implemented as the stacked autoresearch mechanism. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Read an ordered list of phase names, run each phase with its own configuration on the current weights, then evaluate the combined result. The default order is memorize followed by entity-masked SFT: first improve task recall, then apply answer-focused updates.

Implementation and controls

autoresearch/experiment.py recursively dispatches phase configurations and aggregates results. All phases share the outer snapshot bracket. training.stacked.phases in config supplies the order. This differs from summing multiple losses within one gradient step: later phases see weights changed by earlier phases.

Evidence and evaluation

The journal proposes stacking to combine memorize's heldout performance with entity masking's probe behavior. The existence of the configured branch does not establish that the combined recipe inherits both benefits. A final score needs its own measured arm with phase-level telemetry.

Limitations and interpretation

Ordering, cumulative update budget, and interference can change the result. Nested stacked phases or unintended repeated phases need configuration review. Compare the stack with each constituent at a matched budget and re-evaluate after each phase to identify regressions.

Sources

See also