Difficulty-targeted rollout replay
Difficulty-targeted rollout replay is the V-KTO follow-up proposal combining prompt difficulty estimation with reuse of previous rollouts.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Allocate more sampling to prompts that are neither solved nor hopeless, retain useful trajectories in a replay buffer, and reuse them rather than regenerate every candidate. This introduces both a difficulty policy and persistent rollout data, whereas V-KTO samples afresh per step.
Implementation and controls
The June 2 next-arms document ranks this as a candidate and points to difficulty-targeted online selection literature. A reproducible implementation must define difficulty from observed pass rates, the buffer's admission/eviction rules, staleness policy, off-policy treatment, and matched token/wall budgets.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
Old responses may no longer reflect the live policy. A buffer with no coverage or freshness rules can amplify stale errors and leak heldout information. The note's published wall-time claims are not local measurements, and a new loss is not inherently required just to implement a buffer.
Sources
- agi: lile/docs/research/NEXT_ARMS_jun02.md — historical revision
3842fd8875ca. - cont: docs/research/surveys/neurips-2025-sample-efficient.md — checkout audited
87946914c7b9.