TTRL majority-vote training

From The Hei Canon

TTRL majority-vote training is Trainfer's idle-time pseudo-label generator for unlabeled prompts, inspired by test-time reinforcement learning.

Project status: Implemented scheduler, disabled by default; performance/retention gate explicitly deferred. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Collect fresh rollouts for a recent eligible prompt, extract domain-specific equivalent answers, pick a winning answer group, and enqueue SFT on one matching rollout. In code, the vote is plurality with a minimum count of two, not a required strict majority. Ties above the floor break by first occurrence.

Implementation and controls

ttrl_mv.py defines TTRLScheduler, TTRLPolicy, and vote helpers. Defaults: four rollouts, idle threshold 30 seconds, polling every two seconds, three prompts before starting, and at most three uses per inference offset per process. Math keys use answer extraction; code keys use sandbox stdout. Empty stdout does not count as agreement. cfg.ttrl_pseudo_reward defaults false. Sampling occurs outside the training worker; SFT uses the normal queue.

Evidence and evaluation

The module documents smoke coverage and a deferred heldout regression gate. No local TTRL gain on GSM8K is established by that source. The roadmap's published-literature gain is background attribution, not an outcome of this scheduler.

Limitations and interpretation

Agreement is not correctness: repeated hallucinations can win. A two-versus-two tie can be selected deterministically without consensus. A verifier claiming a prompt or extracting an answer is weaker than externally verifying the answer. Per-offset caps reset on restart. The docstring's SFT safety claim is limited by the same shared-parameter/optimizer issues as other SFT recipes.

Sources

See also