Trainfer learning methods
Trainfer learning methods is the source-audited catalogue of methods introduced, implemented, tested or concretely proposed in agi/lile, Trainfer and Cont. It covers the current objective registry, every training-mechanism branch in cont's main autoresearch runner, the recovered May–June 2026 experimental arms, the method tier menu and roadmap candidates. Trainfer research literature preserves the complete project survey accounts for further paper-only methods.
Audit date: 14 September 2026. This catalogue documents research; it does not establish that continual learning has been solved. Each method page separates mechanism, implementation/controls, measured evidence, and limitations. “Implemented” means found in the audited source, not necessarily enabled in a running daemon; “proposed” does not mean shipped. Standard methods are adaptations, not claims of project invention.
Reading the evidence
- V-KTO means verifier-graded KTO with adaptive rollouts; its June pilot regressed severely after 51 updates. The implementation survives in agi history.
- CCPD and CCD are different objectives: critique-based candidate ranking versus context teacher/student matching.
- KTO is the loss; V-KTO and the binary-verifier Arm 3 are data/sampling recipes using it.
- The early GSM8K five-example fine-tune reached 44%, but matched five-shot ICL reached 96%. The journal withdrew its SOTA claim.
- “Razin-safe” in older source comments is not a certificate for actual AdamW/LoRA behavior. See Razin safety and Razin safety monitor.
- Reported numbers belong to their model, suite, seed and date. Pending table cells in old templates are not completed experiments; prose hypotheses are not measured mechanisms.
Implemented objective primitives
- Supervised fine-tuning (Trainfer) — Registered as
sftandweighted_sft. - Next-token prediction (Trainfer) — Registered as
ntp. - KTO — Registered as
kto; also used by historical verifier-driven experiments. - Chain of Hindsight — Registered as
coh; pure-CoH historical ablation was pre-registered. - Hinge contrastive SFT — Registered as
hinge. - CCPD — CCPD v2 is conditionally registered as
ccpd_v2; v1 is a superseded design. - Context distillation (Trainfer) — Registered as
ccd; “fact-preserving” is its design goal, not a proven guarantee. - KL anchor — Registered as
kl_anchor, including prompt, full-sequence, and target-position scopes. - Surgical unlikelihood — Registered as
unlike. - Razin safety monitor — Registered as
safety_monitor; observational sidecar, not a blocking safety gate.
Learning and replay recipes
- Greedy memorization — Implemented in
memorize.pyand used by the autoresearch runner. - Entity teaching — Implemented as
teach_entityandteach_number. - Trace infilling — Prefix masking is implemented; the plan's richer critique-targeted multi-candidate infilling is a broader design.
- RLVR combined-loss training — Standalone scheduler implemented; the generic autoresearch runner's
rlvrbranch is still a placeholder. - Pre-sampled unlikelihood training — Implemented as
hybrid_presample_unlikein the autoresearch runner. - Self-synthesized paraphrase training — Implemented as
self_synthin the autoresearch runner. - Entity-masked SFT — Implemented as
entity_masked_sft. - Stacked training recipes — Implemented as the
stackedautoresearch mechanism. - TTRL majority-vote training — Implemented scheduler, disabled by default; performance/retention gate explicitly deferred.
- Idle feedback replay — Implemented scheduler, controlled by
idle_replay(default false).
Historical experimental arms
- V-KTO — Implemented in historical agi code; the June 2026 pilot was abandoned after severe regression.
- Rejection fine-tuning — Implemented in historical
experiment_pi_old_synth.py; tested on Track B. - STaR-style canonical fallback — Implemented as historical
--mode star; tested on Track B.
Optimization, state and evaluation
- Per-objective optimizer isolation — Implemented optional mode; default off.
- Lion8bit optimization — Selection implemented as
optimizer="lion8bit"; not the default. - Progressive LoRA residual consolidation — Implemented in
ModelState.merge_active_into_residual. - Frozen-reference anchoring — Implemented and opt-in through
cfg.frozen_ref. - Autoresearch recipe optimization — Implemented workflow; individual configured branches have different implementation/completion states.
- Snapshot-bracket evaluation — Required by the research contract; implementation and historical failures are documented.
- In-context learning baseline — Measured baseline in the logical, GSM8K, and HumanEval campaigns.
Concrete proposals and comparison methods
- AdaSTaR-RFT — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Intruder-dimension damping — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- FLOW weighting — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- EOS preference regularization — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Active preference-query selection — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Pretraining replay injection — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Low-confidence replay weighting — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Hypernetwork LoRA generation — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Streaming DataInf — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- SEAL-style self-editing — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Difficulty-targeted rollout replay — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Hint-RFT — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Function-vector anchoring — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- SCoRe-style correction — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Rehearsal and snapshot self-distillation — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Deferred feedback batching — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
- Optimizer research candidates — Mixed proposal/deferred status; none of the alternatives below is a registered choice in the audited Trainfer optimizer selector.
- RFT research variants — Literature/design candidates, distinct from the measured local RFT, STaR, KTO and V-KTO arms.
- DPO and GRPO in Trainfer research — Background/comparison methods; no
dpo,ipo,ppo, orgrpokey in the audited objective registry.
Coverage and provenance
The objective coverage is sft, weighted_sft, ntp, kto, coh, hinge, kl_anchor, safety_monitor, unlike, ccd and conditionally ccpd_v2. SFT and weighted SFT share an article; trace infilling is a mask capability rather than a registry key. DPO/IPO/PPO/GRPO are discussed in the plan but are not keys in this registry.
Current code sources are pinned to their audited revisions. Historical articles cite agi revision 3842fd8; locally that history was reachable through origin/feat/lion8bit-optimizer. A branch name does not establish which optimizer an experiment used. A reproducible source archive, SHA-256 manifest, wikitext pages, and seed script live in ~/ht/wiki/learning_methods_2026_09_14/ and ~/ht/wiki/seed_learning_methods_2026_09_14.py.
Sources
- trainfer: trainfer/objectives/__init__.py — checkout audited
1c6391f3773b. - trainfer: trainfer/PLAN.md — checkout audited
1c6391f3773b. - cont: autoresearch/config.json — checkout audited
87946914c7b9. - cont: autoresearch/experiment.py — checkout audited
87946914c7b9. - cont: docs/research/JOURNAL.md — checkout audited
87946914c7b9. - cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - agi: autoresearch/LATTICE.md — historical revision
3842fd8875ca.