Trainfer research survey/iclr-2025-2026
ICLR 2025 / 2026 — test-time adaptation, continual LoRA, sparse-feedback RL is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.
Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.
ICLR 2025 / 2026 — test-time adaptation, continual LoRA, sparse-feedback RL
- Scope: ICLR 2026 accepted + ICLR 2025 load-bearing; filtered against existing
trainfercitations. - Compiled: 2026-04-17 (background agent
a9913f74c9375c45e).
Tags: [STRONG] · [RELEVANT] · [BACKGROUND] · [SKIP].
---
ICLR 2026 — highest priority
1. In-Place Test-Time Training (Oral) [STRONG]
- <https://openreview.net/forum?id=dTWfCLSoyl>
- Reuses the final MLP projection matrix as fast-weights and aligns the TTT objective with NTP so adaptation is drop-in for any LLM. Processes large chunks in parallel so inference-time updates stay cheap.
- Why it matters: strongest ICLR 2026 argument that live-weight updates at serving time are feasible and architecture-compatible — direct theoretical cover for
trainfer.
2. Reward Is Enough: LLMs Are In-Context RL Learners [RELEVANT]
- <https://openreview.net/forum?id=keCXNHOe4W>
- LLMs convert scalar reward feedback at inference into RL-like self-improvement via multi-round prompting (ICRL), outperforming test-time scaling baselines.
- Why it matters: cite-magnet framing for "sparse scalar feedback is trainable signal" that matches
trainfer's user-feedback loop exactly.
3. Meta-UCF — Unified Task-Conditioned LoRA Generation for Continual Learning [STRONG]
- <https://openreview.net/forum?id=iNg5KL7eTC>
- Hypernetwork ingests a task-embedding and emits layer-wise rank-r LoRA deltas with meta-contrastive orthogonality; memory-constant continual adaptation over unbounded task streams (+2.2pp avg, −13% forgetting).
- Why it matters: per-user/per-task LoRA generation without inner-loop gradients — natural mode for
trainfer's identity-LoRA direction.
4. Doc-to-LoRA — Instantly Internalize Contexts [STRONG]
- <https://openreview.net/forum?id=r9NMVtrJGj>
- Hypernetwork meta-learns to produce a LoRA adapter from a prompt/document in a single forward pass.
- Why it matters: closest 2026 analog to "materialize a user-LoRA on the fly" — contrast point for
trainfer's live training.
5. Uni-DPO — Unified Dynamic Preference Optimization [RELEVANT]
- Repo. Jointly weights preference pairs by intrinsic quality and current model fit; focal-style down-weighting of well-fitted pairs reduces overfit; beats DPO/SimPO across text/math/multimodal.
- Why it matters: drop-in upgrade over KL-regularized DPO for
trainfer's sparse-feedback regime.
6. Data Selection for Efficient Preference Alignment [RELEVANT]
- arXiv:2508.07638. ~30% of preference data, chosen well, matches full-data DPO.
- Why it matters: justifies that
trainfergets a handful of signals per user and that's sufficient.
7. ∇-Reasoner — Test-Time Gradient Descent in Latent Space [BACKGROUND]
- <https://openreview.net/forum?id=pEJAja73dk>
- DTO refines logits with gradients from likelihood + reward inside the decode loop; proves duality to KL-regularized RL.
- Useful framing for "inference-time gradient updates = implicit RL", relevant for positioning but not the LoRA path.
- <https://openreview.net/forum?id=tWAnCRYMcT>
- Maps when test-time adaptation (context, memory, weights) helps vs. hurts under verifiable feedback; explicit accuracy-vs-compute Pareto curves.
- Why it matters: exactly the evaluation frame
trainferneeds — position the daemon's modes on this Pareto.
9. Aligner, Diagnose Thyself — Meta-Learning for Intrinsic Feedback [RELEVANT]
- ICLR 2026 Downloads.
- Meta-learns how to combine the model's own introspective feedback signals during alignment.
- Why it matters: closest 2026 descendant of Chain-of-Hindsight — contributes to the "learn from NL critique" axis.
10. Alignment through Meta-Weighted Online Sampling [RELEVANT]
- ICLR 2026 Downloads.
- Bridges online data generation and preference optimization via meta-weighted sampling.
11. In-Context Adaptation (ICA) [BACKGROUND]
- <https://openreview.net/forum?id=f58uDOwLaq>
- Framework for OOD generalization via in-context adaptation rather than weight updates.
- Foil to
trainfer— argues the training-free axis; worth citing to defend why weight-space adaptation is still needed.
ICLR 2025 — load-bearing keepers
12. Spurious Forgetting in Continual Learning of Language Models [STRONG]
- <https://openreview.net/forum?id=ScI7IlKGdI>
- Continual-learning drops often reflect task-alignment collapse (orthogonal weight updates), not true knowledge loss; simple bottom-layer freezing strategy recovers most of the loss.
- Why it matters: directly informs
trainfer's layer-selection and regularization; reframes what "forgetting" even means.
13. SD-LoRA — Scalable Decoupled LoRA for Class Incremental Learning [RELEVANT]
- Decouples LoRA components to scale incrementally across tasks without rehearsal.
- Why it matters: architectural template for stacking / rotating adapters as
trainferaccumulates user sessions.
14. Unlocking Function Vectors for Catastrophic Forgetting in Continual Instruction Tuning [STRONG]
- Function-vector-guided KL regularizer curbs forgetting during continual instruction tuning.
- Why it matters:
trainferalready has KL infrastructure — this is a ready upgrade path, and complements Spurious Forgetting nicely.
15. HMoRA — Hierarchical Mixture of LoRA Experts [RELEVANT]
- <https://iclr.cc/virtual/2025/poster/28518>
- Hierarchical routing over LoRA experts with balance/certainty auxiliary losses.
- Why it matters: natural scaffold for per-user identity-LoRA mixed into a shared expert pool.
Skipped
SELF-CRITEACH / Harness-as-Policy (workshop-only); GoodPoint (review-domain); Active-Critic (NLG eval only); BA-LoRA, Co-LoRA (tangential). XPO / COPO are 2025-era and Async-RLHF neighborhood already covers that ground.
---
Sources
Archived source
- cont: docs/research/surveys/iclr-2025-2026.md — checkout audited
87946914c7b9.