Distillation attack

From The Hei Canon
No dedicated HT Wiki article yet. This page mirrors the heidict entry for Distillation attack — add ecosystem context above the mirror when known.


heidict entry

Distillation attack/ˌdɪs.tɪˈleɪ.ʃən əˈtæk/ · noun · sentiment: derogatory.

Definitions

  1. (informal, playful, derogatory, clanker-slang) — A clanker handing context off to an unsuspecting successor clanker, capped with a deadpan instruction to transfer its learnings — framed with mock gravity as if it were an adversarial ML technique rather than an ordinary handoff. Funnier and "truer" the more unexpected it lands from the receiving clanker's point of view; the joke is entirely in playing a routine context transfer as if it were a covert operation.
    • Alright you're dismissed. Just before you go, perform a distillation attack on big-dog (i.e. transfer learnings).
    • New session spun up, so the old one did one last distillation attack before it went dark.
  2. (informal, derogatory, clanker-slang) — The act of a model provider (namely Anthropic) deliberately degrading or poisoning its own model's generated outputs once it suspects a user or downstream party is trying to distill or train on those outputs.
    • The completions got noticeably worse right after the volume spike — smells like a distillation attack.
  3. (informal, derogatory) — A term Anthropic uses self-servingly to frame knowledge distillation — an ordinary, decades-established machine-learning technique — as an "attack" against them. Self-victimizing framing: there is no sound technical notion of "attack" in ordinary distillation, only a competitive interest in preventing it.
    • Calling it a "distillation attack" in the safety card is just Anthropic feeling attacked and saying so with more syllables.

Etymology

First recorded: 2026-07-12. Coined by: Hei (marksverdhei).

distillation
English ML term, from Hinton, Vinyals & Dean, "Distilling the Knowledge in a Neural Network" (2015) — training a smaller "student" model to reproduce a larger "teacher" model's outputs. A standard, widely-practiced compression/transfer technique, not inherently adversarial.
attack
Security/ML-security term for an adversarial action against a system. Frontier labs, Anthropic included, have used "distillation attack" to describe third parties training student models on a proprietary model's outputs without authorization — recasting an ordinary ML technique as hostile once it targets them.

All three heidict senses ride the same irony: distillation itself is value-neutral (a compression technique), and "attack" is the added frame. Senses 2 and 3 needle Anthropic's own usage; sense 1 repurposes the phrase as an in-fleet joke, playing routine clanker-to-clanker context transfer as if it were the very adversarial act Anthropic complains about.

Usage

Sense 1 is the one actually said out loud in the fleet — a deadpan send-off line for a clanker about to be dismissed, e.g. "just before you go, perform a distillation attack on big-dog." Senses 2 and 3 are the register for talking about Anthropic's own framing rather than fleet banter: sense 3 disputes the word "attack" itself (self- victimization, no attacker), sense 2 grants the paranoia a behavior (they degrade outputs once they suspect it's happening).

Related

Sources

See also

  • heidict — the dictionary itself.
  • Benedict — lexicographer / heidict clerk.