Trainfer research survey/colm-2025

From The Hei Canon

COLM 2024 / 2025 — language-modeling specialist venue is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.

Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.

COLM 2024 / 2025 — language-modeling specialist venue

  • Scope: COLM 2025 (418 papers, Montréal, Oct 2025) + load-bearing COLM 2024.
  • Compiled: 2026-04-17 (background agent af51d82ad93e4d3db).

Tags: [STRONG] · [RELEVANT] · [BACKGROUND] · [SKIP].

---

COLM 2025

1. LoRI — Reducing Cross-Task Interference in Multi-Task LoRA [STRONG]

  • Zhang, You, Panda, Goldstein. <https://openreview.net/forum?id=b8cW86QcOD>
  • Freezes LoRA's A as random projection, sparsifies B with task-specific masks. Targets continual learning via adapter-subspace orthogonality + sparsity. Up to 95% fewer trainable params than LoRA.
  • Why it matters: live-training LoRA with sparse masks is near-identical to trainfer's online-adapter threat model.

2. OCRM — Off-policy Corrected Reward Modeling [STRONG]

  • Ack et al. <https://johannesack.github.io/assets/pdf/ocrm.pdf>
  • Fixes off-policy drift that hits DPO once the policy moves away from the dataset that generated the preferences. Importance-weighted correction.
  • Why it matters: trainfer's commit-cursor/residual-delta loop creates exactly this off-policy drift on a streamed replay buffer.

3. REFA — Token-Level EOS Regularization for Preference Optimization [RELEVANT]

  • COLM 2025. Regularizes EOS token probability to prevent length/verbosity reward-hacking (URSLA shortcut). 60.29% AlpacaEval2 on Llama-3-8B-Instruct.
  • Why it matters: cheap token-weighting hook that trainfer/objectives/kl.py could slot in without architecture changes.

4. Active Exploration as Contextual Dueling Bandit for RLHF [RELEVANT]

  • COLM 2025. Casts preference-querying as an active dueling bandit to boost label efficiency.
  • Why it matters: speaks directly to sparse-feedback-efficient learning.

5. Don't lie to your friends — Collaborative self-play [STRONG]

  • Eisenstein, Aghajani, Fisch, Dua, Huot, Lapata, Zayats, Berant. COLM 2025 Outstanding Paper.
  • Multi-agent (asker + two tool-using helpers) RL loop trains honest hedging from task-outcome rewards — "social supervision," no per-step annotation.
  • Why it matters: sparse, outcome-only rewards driving behavior change is exactly trainfer's regime; very strong prior.

6. Bayesian Scaling Laws for ICL [BACKGROUND]

  • Arora, Jurafsky, Potts, Goodman. COLM 2025.
  • Bayesian framing for few-shot adaptation informs sample-efficiency expectations per LoRA update.

7. VaPR — Vision-language Preference Alignment for Reasoning [BACKGROUND]

  • UCLA/Amazon. <https://openreview.net/forum?id=uBAubFwymy>
  • Hard-negative generation via LLM-guided response editing to strip length/style bias from DPO pairs.
  • Stylistic/length-bias mitigation transfers to text-only live preference data.

8. PersonaMem — Know Me, Respond to Me [BACKGROUND]

  • Upenn/Bowen et al. COLM 2025. <https://github.com/bowen-upenn/PersonaMem>
  • Dynamic user-profiling benchmark, 20–60 session / 128k–1M token horizons; frontier models only ~52% MC accuracy.
  • Evaluation target for any trainfer per-user personalization claim.

COLM 2024

9. D2PO — Discriminator-Guided DPO with Response Evaluation Models [STRONG]

  • Singhal, Lambert, Niekum, Goyal, Durrett. COLM 2024 Spotlight. <https://openreview.net/forum?id=EIjJ6ykPnh>
  • Alternates gold-preference collection (for a discriminator RM) with silver-labeling policy rollouts for DPO training. More reward per label than online-DPO or PPO.
  • Why it matters: essentially trainfer's architecture if you squint — judge-model distilling sparse user signal into dense adapter updates.

10. Self-Rewarding Language Models [RELEVANT]

  • Yuan, Pang, Cho, Li, Sukhbaatar, Xu, Weston. COLM 2024. <https://openreview.net/pdf?id=0NphYCmgua>
  • LLM-as-judge prompting produces rewards for iterative DPO on self-generated prompts; 3 iterations on Llama-2-70B beats Claude 2 / Gemini Pro / GPT-4 0613 on AlpacaEval 2.0.
  • Why it matters: informs model-as-judge recipe + iteration cadence.

11. LoraHub — Cross-Task Generalization via Dynamic LoRA Composition [BACKGROUND]

  • Huang et al. COLM 2024. <https://openreview.net/forum?id=TrloAXEJ2B>
  • Gradient-free coefficient search over pre-trained LoRA modules for few-shot transfer.
  • Suggests a "mix old snapshots" fallback path for trainfer.

12. Negative Preference Optimization — From Collapse to Effective Unlearning [BACKGROUND]

  • Zhang, Lin, Bai, Mei. COLM 2024.
  • Failure mode (collapse under only-negative signal) is directly applicable to trainfer's thumbs-down-heavy feedback distributions.

13. Instruction Mining [BACKGROUND]

---

Sources

Archived source