RLVR combined-loss training

From The Hei Canon

RLVR combined-loss training is cont's rollout–judge–train loop for reinforcement-learning-style training with verifiable or judged feedback.

Project status: Standalone scheduler implemented; the generic autoresearch runner's rlvr branch is still a placeholder. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Sample candidate responses, obtain grades plus optional critiques/counterfactuals, then build a weighted sum of existing primitives. Correct responses route to weighted SFT; wrong responses with critique can route to CoH; ambiguous responses route to undesirable KTO; tokenized wrong/counterfactual corrections can route to unlike. A KL sidecar anchors the update.

Implementation and controls

cont/teach/rlvr_loop.py provides RLVRConfig, build_combined_spec, and RLVRScheduler. Defaults include K=4, temperature 0.8, top_p=0.95, and maximum 512 new tokens. Default coefficients are SFT 0.1, CoH 1, KTO 1, unlike 0.5, KL 0.05. Missing tokenizer or critique can suppress routes, so log emitted objective counts, skipped reasons, and commit tokens. Teacher modules supply structured grades; a teacher judgement is not automatically a ground-truth verifier. The separate historical HumanEval runner constructs binary grades from EvalPlus. Current autoresearch/experiment.py explicitly returns “rlvr arm not wired yet”; choosing that menu string does not invoke the standalone scheduler.

Evidence and evaluation

C-002 scoped a seven-arm mixture comparison. Historical HumanEval reports separate buggy routing (mean −6.8 points), corrected no-unlike routing (−19.3 points), and unlike-enabled seed 0 (−18.8 points). Silent dropped objectives make the buggy arm a different effective method, not a legitimate equivalent control.

Limitations and interpretation

“RLVR” here names an orchestration recipe; it does not imply PPO or GRPO is the actual loss. Grade distribution, missing-field routing, canonical answers, and sampling configuration determine the update. Compare actual submitted specs, tokens and task accuracy, not just requested weights. No general gain can be inferred from successful API submissions.

Sources

See also