SCoRe-style correction

From The Hei Canon

SCoRe-style correction is the plan's multi-turn self-correction training tier.

Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Treat an initial response and a subsequent correction as a trajectory. Train toward improved later responses using judged or verifiable improvement, rather than treating the first answer as an isolated labelled sample.

Implementation and controls

Tier T3.2 in the plan sketches multi-turn GRPO-style training and possible critique-conditioned or user-supplied targets. The audited objective registry has no SCoRe or GRPO primitive. A runnable adaptation needs trajectory sampling, advantage definition, masks, stopping rules, and a correction-specific benchmark.

Evidence and evaluation

The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.

Limitations and interpretation

Ordinary CoH text formatting or a second chat turn is not full SCoRe training. A model can revise correct answers into wrong ones, so measure both correction of mistakes and preservation of initially correct responses. The plan's published benchmark gains do not transfer automatically to open-ended local chat.

Sources

See also