Rejection fine-tuning

From The Hei Canon

Rejection fine-tuning is the historical best-of-N verifier-selected SFT arm, also described as RFT or rejection-sampling SFT.

Project status: Implemented in historical experiment_pi_old_synth.py; tested on Track B. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Generate responses to a task, verify each extracted answer, and select the shortest passing response, breaking ties by original position. Skip groups with no passing responses and groups where all responses pass. Train SFT on the chosen response. The all-pass skip was intended to avoid redundant self-distillation.

Implementation and controls

The historical runner's filter_and_rank(mode="rft") performs selection, with collection separated from later training. Passing is task-verifier agreement, not a language-model critique score. Training uses the existing SFT API rather than a new objective. Generated rollouts, selection decisions, skipped reasons, and post-evaluations form the reproducibility record.

Evidence and evaluation

The Track B Qwen3-0.6B lattice records Arm 1 at −11.25 percentage points versus cold. Its controls include canonical-fallback STaR and binary KTO on the same mechanism family experiment. This result is specific to that corpus/model and training setup.

Limitations and interpretation

Selecting a shortest response is a surface-length preference that can remove useful reasoning. Verifier errors or extraction quirks contaminate the target set. Skipping extremes changes the effective sample budget, so nominal prompt count is not actual training count. The local negative result is not evidence that all published RFT variants fail.

Sources

See also