Difficulty-targeted rollout replay

From The Hei Canon

Difficulty-targeted rollout replay is the V-KTO follow-up proposal combining prompt difficulty estimation with reuse of previous rollouts.

Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Allocate more sampling to prompts that are neither solved nor hopeless, retain useful trajectories in a replay buffer, and reuse them rather than regenerate every candidate. This introduces both a difficulty policy and persistent rollout data, whereas V-KTO samples afresh per step.

Implementation and controls

The June 2 next-arms document ranks this as a candidate and points to difficulty-targeted online selection literature. A reproducible implementation must define difficulty from observed pass rates, the buffer's admission/eviction rules, staleness policy, off-policy treatment, and matched token/wall budgets.

Evidence and evaluation

The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.

Limitations and interpretation

Old responses may no longer reflect the live policy. A buffer with no coverage or freshness rules can amplify stale errors and leak heldout information. The note's published wall-time claims are not local measurements, and a new loss is not inherently required just to implement a buffer.

Sources

See also