Low-confidence replay weighting

From The Hei Canon

Low-confidence replay weighting is the proposal to prioritize uncertain examples for additional learning.

Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Use the policy's average response log probability as a cheap confidence proxy and raise replay priority for low-confidence events. This aims to spend updates on informative, difficult examples instead of only repeating easy ones.

Implementation and controls

Roadmap PR O proposes a twofold replay-priority multiplier and a 500-event mixed-feedback comparison. Its gate asks for gains on at least two harness tasks without a measurable queue-latency regression. This is a scheduling/weighting change, not a new objective.

Evidence and evaluation

The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.

Limitations and interpretation

Low confidence is not synonymous with correct, difficult, or informative: it can flag noise and errors. Response length normalization and cross-objective calibration matter. It pulls in a different direction from FLOW's preference for base-familiar examples; neither should be assumed superior without matched testing.

Sources

See also