Hinge contrastive SFT

From The Hei Canon

Hinge contrastive SFT combines positive-response SFT with a clipped preference-margin penalty.

Project status: Registered as hinge. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Let c and r be mean per-token log probabilities of chosen and rejected responses. The loss is −mean(c) + rejection_weight × mean(max(0, r − c + margin)). The hinge term vanishes once the chosen response clears the margin; positive SFT remains.

Implementation and controls

objectives/hinge.py::hinge_contrastive_loss consumes prompt, chosen, and rejected. Defaults: margin 1.0 and rejection weight 0.5. Both responses are chat-templated and length-normalized separately. No auxiliary generation or explicit reference model is required.

Evidence and evaluation

The CCPD ranking report records a “hinge primary” deployment decision at moderate ranking reliability. This describes a fallback preference in the design/status record; it does not mean every current feedback route invokes hinge. The audited replay router maps preferred rewrites to weighted SFT.

Limitations and interpretation

Clipping bounds the active margin region, not the total influence of repeated updates. Its negative-response term means the full objective does not inherit the positive-only SFT theorem. Margin satisfaction and preference correctness are separate questions; response-level mean log probabilities can conceal local errors.

Sources

See also