Trainfer research survey/icml-2025

From The Hei Canon

ICML 2025 — sample-efficient LLM learning is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.

Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.

ICML 2025 — sample-efficient LLM learning

  • Scope: ICML 2025 accepted, filtered against existing trainfer citations.
  • Compiled: 2026-04-17 (background agent a540a3fd08201d78e).

Tags: [STRONG] · [RELEVANT] · [BACKGROUND] · [SKIP].

---

Tier 1 — Directly load-bearing

1. FLOW — Upweighting Easy Samples Mitigates Forgetting [STRONG]

  • Sanyal, Prairie, Das, Kavis, Sanghavi. arXiv:2502.02797. poster · OpenReview.
  • Parameter-free sample-weighting using the pre-trained model's own per-example loss: up-weight samples where the base model already scores well. Operates in sample space rather than gradient space, so stacks cleanly with CCPD.
  • Why it matters: drop-in weight for SFT/NTP heads with no extra compute; especially useful in the 10–100 req/hr regime where every sample carries high variance.

2. ConfPO — Policy Confidence for Critical Token Selection [STRONG]

  • Yoon et al. arXiv:2506.08712. poster.
  • Token-level DPO variant that restricts gradient to tokens where the policy is confident; tighter KL-budget use, less reward hacking, no auxiliary model.
  • Why it matters: lets KTO/hinge heads concentrate gradient on preference-critical tokens without an extra scorer — aligns with kl_scope.

3. EXPO — Explicit Preference Optimization [RELEVANT]

  • Hu, Kong, He, Wipf. poster.
  • Proves DPO-style reparameterized objectives inherit sub-optimal regularization; proposes an explicit preference loss that avoids the implicit-reward pathology.
  • Why it matters: candidate replacement/ablation for the KTO head.

4. Theoretical Analysis of KL-regularized RLHF with Multiple Reference Models [STRONG]

  • arXiv:2502.01203. page.
  • First exact solution for reverse-KL RLHF with several anchors; O(1/n) sub-optimality and O(1/√n) optimality sample complexity; validated on Qwen 2.5 with GRPO on GSM8K.
  • Why it matters: trainfer/objectives/kl.py anchors to one reference — this paper gives the math for anchoring to base + snapshot + adapter, matching the live-training snapshot model.

5. PF-PPO — Policy Filtration for RLHF [RELEVANT]

  • poster. RM reliability varies with reward magnitude → filter unreliable samples before PPO update.
  • Why it matters: maps cleanly onto trainfer's replay-queue gating under noisy reward.

6. PILAF — Policy-Interpolated Learning for Aligned Feedback [STRONG]

  • poster. Active sampling method for preference queries that interpolates between policy and reference.
  • Why it matters: 10–100 req/hr = the exact active-query regime PILAF is designed for — decides which prompts to solicit feedback on.

Tier 2 — Worth citing / implementing

7. TLM — Test-Time Learning for LLMs [RELEVANT]

  • Fhujinwu et al. poster · code.
  • Unlabeled test-data adaptation via input-perplexity minimization, high-perplexity prioritization + LoRA for stability.
  • Why it matters: perplexity-based active-selection is transplantable as a no-feedback-required auxiliary objective.

8. Surprising Effectiveness of Test-Time Training for Few-Shot Learning [BACKGROUND]

  • poster. 6× ARC and +7pp BBH gains from per-example TTT with tiny LoRA updates.
  • Empirical validation that micro-LoRA test-time updates — trainfer's core design — work on hard reasoning tasks.

9. Flat-LoRA — LoRA over a Flat Loss Landscape [STRONG]

  • OpenReview. Seeks flat minima in the *full* parameter space from the LoRA manifold using a Bayesian-expectation loss, avoiding SAM's double-backprop cost.
  • Why it matters: flatness-aware LoRA update rule with no added forward-pass — candidate default for the snapshot optimizer.

10. TSAM — Tilted Sharpness-Aware Minimization [BACKGROUND]

  • poster. Exponential-tilt smoothing of SAM's min-max; easier to optimize, flatter minima than SAM.
  • Fallback if Flat-LoRA's Bayesian surrogate underperforms.

11. Scaling Laws for Forgetting with Pretraining Data Injection [STRONG]

  • Apple, ICML 2025. poster. Injecting as little as 1% pre-training data prevents forgetting across scales/domains.
  • Why it matters: concrete floor for trainfer's replay mixture ratio — plug-in number for replay_streams.

12. From RAG to Memory + Cut-and-Replay [BACKGROUND]

  • arXiv:2502.14802 + companion. Non-parametric continual learning alternatives.
  • Fallback path if the parametric daemon ever needs a complement.

Skipped (off-scope)

InfAlign (inference-time alignment), Data Mixing Optimization for SFT (pre-training scale), MM-RLHF (multimodal), LowRA (sub-2-bit quant), ProLoRA (diffusion), MetaOptimize (general LR tuner), Reward-Model Robustness papers, DOOR/W-DOOR (safety refusal).

---

Synthesis

Highest-leverage picks for immediate citation: FLOW, ConfPO, Multi-Reference KL-RLHF, PILAF, Flat-LoRA, TLM, Apple's 1%-injection scaling law. FLOW and the 1%-injection law are the two that could change trainfer code this sprint.

Archived source