DPO and GRPO in Trainfer research

From The Hei Canon

DPO and GRPO in Trainfer research explains how preference and policy-gradient methods appear in the project's design vocabulary without implying that every named algorithm ships as a primitive.

Project status: Background/comparison methods; no dpo, ipo, ppo, or grpo key in the audited objective registry. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

The plan compares paired preference objectives such as DPO/IPO with explicit likelihood training and critique-driven policy gradients. Its CCPD revision separates detached critique scoring from the trainable policy likelihood. GRPO-style group-relative training appears in the multi-turn correction design, while implemented CCPD uses its own centered ranks and REINFORCE-style term.

Implementation and controls

Use the registry, not a preferred/rejected API shape, to identify the actual loss. The shipped primitives include KTO, hinge, SFT-family objectives, CCD and optional CCPD v2. A DPO-shaped input can route into hinge, weighted SFT, or CCPD, and RLVR orchestration composes those primitives. No separate reward model, PPO clipping loop, or faithful GRPO implementation follows merely from those labels.

Evidence and evaluation

The plan and surveys cite these methods as alternatives and motivate the likelihood-displacement analysis. The historical lattice measures specific local arms; it does not establish a DPO/IPO/GRPO head-to-head benchmark.

Limitations and interpretation

The absence claim is about the audited project registry, not all upstream Unsloth/TRL capabilities. Comparisons of safety must use the actual loss and optimizer, not the plan's broad taxonomy. Newly introduced objectives should receive their own implementation and evidence entry if they later land.

Sources

See also