V-KTO

From The Hei Canon

V-KTO is the historical verifier-graded KTO recipe with adaptive rollouts developed in agi/lile, combining verifier-based response ranking with the existing KTO loss. It is a project experiment, not a separate registered daemon objective.

Project status: Implemented in historical agi code; the June 2026 pilot was abandoned after severe regression. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

For each HumanEval prompt:

  1. Generate eight candidates at temperature 0.8.
  2. Compute fractional correctness f = (base tests passed + EvalPlus tests passed) / (base tests + EvalPlus tests).
  3. If max(f) − min(f) < 0.1, generate additional candidates to reach 16 and then 32; additional samples use temperature 1.0.
  4. If the spread is still below 0.1, skip the prompt.
  5. Sort by (score, original index). Label the top two desirable and bottom two undesirable; discard the middle.
  6. Train KTO at objective weight 1.0 plus a target-position KL anchor at weight 0.05.

The fractions select labels; they do not become continuous rewards or loss weights. KTO receives four independently labelled examples rather than a DPO-style pairwise margin.

Implementation and controls

Historical autoresearch/experiment_humaneval.py implements _vkto_step and _build_vkto_spec; --mode vkto forces the KTO-only profile. The builder bypasses build_combined_spec, despite the early scope document naming that helper. verify_fractional is in the historical HumanEval verifier; evaluation remains binary all-tests-pass. The smoke configuration used seed 0, learning rate 0.0002, max 2,000 generated tokens, and thinking disabled. Adaptive K is separate from the generic k_rollouts=4 metadata. Failed generation calls can leave fewer than the nominal tier count. A default 200-step cap and 20-entry correctness window at threshold 0.95 stop the training loop; skipped prompts do not enter that window. Introduction commits: 780c57c (scope), 7b9ef5e (fractional verifier), ab134ce (runner), all 24 May 2026. The implementation was found in history, reachable through the locally available origin/feat/lion8bit-optimizer ref; it is absent from the audited main checkouts.

Evidence and evaluation

The 1 June smoke performed 20 iterations and 13 submissions in 4,401 seconds; post-evaluation was 20/64 (31.2%). It was designated a plumbing check. The longer Qwen3-1.7B 4-bit run stalled at 51 training submissions. The 2 June salvage pilot reported:

Task N Cold V-KTO Change
HumanEval heldout 64 34.4% 1.6% −32.8 percentage points
MBPP 30 50.0% 0.0% −50.0 points
GSM8K 30 46.7% 36.7% −10.0 points
LogicBench 20 40.0% 25.0% −15.0 points

The report records token-repetition loops and interprets the submission plateau as uniformly collapsed outputs failing the spread filter. The planned multi-seed campaign became a single-seed go/no-go pilot without the full ICL/SFT falsification suite.

Limitations and interpretation

Relative labels need not represent absolute correctness: the smoke log includes bottom scores [0.0, 1.0], so a fully passing response was labelled undesirable. A partly correct top response can likewise be desirable. There is no explicit tie-group rejection or candidate deduplication in this selection code. The observed failure does not isolate adaptive K, fractional ranking, learning rate, label construction, or anchoring as the cause. The report's claim that an asymmetric cold/post contingency table alone rules out a snapshot bug is stronger than those counts establish; restoration needs independent state and inference checks. Its one-sided p-values test improvement, not statistical significance of damage. No general theorem that KTO or sharper verifier signals necessarily collapse follows from this pilot.

Sources

See also