Model Distillation Ethics

Debate over whether distillation constitutes theft or legitimate technique, comparison to how GPT trained on internet data, noting API fees were paid, irony of Anthropic criticizing distillation while pirating books for training data

← Back to Notes on DeepSeek

The debate over model distillation reveals a sharp divide between those who view it as "theft" and those who see it as a natural, recursive progression of technology akin to how GPT distilled the internet. Commenters highlight the irony of major AI labs condemning the practice while facing their own legal battles over pirated training data, noting that distillation at least involves paying for API access. While some argue that relying on distillation prevents a developer from ever reaching the true technological frontier, others defend it as a pragmatic strategy for efficiency, comparing it to the common use of generic brands in other industries.

7 comments tagged with this topic

View on HN · Topics
You wouldn’t steal a brain
View on HN · Topics
What's wrong with distillation? Wasn't GPT a distillation of the world's internet? That's how technology levels proceed, by recursively consuming the previous ones.
View on HN · Topics
It's absolutely mind boggling to see claims of model distillation being theft, a class of attack, and all sorts of claims all the while Meta is in court for copyright violation, anthropic has had to settle a case with authors. With distillation "attacks" at least they paid API fees.
View on HN · Topics
Anthropic had to settle with authors because they literally pirated books! Their behavior regarding distillation is genuinely beyond parody.
View on HN · Topics
There are 2 things worth separating. 1) China distills and is therefore morally bad. As you rightly point out, that's not a great argument. 2) China distills and is therefore possibly not that competent. I think that makes sense. If they only catch up to the frontier through distillation then 1) Their model will never be as good as the model they are distilling from. 2) They will never reach the frontier - they need someone else to do it first.
View on HN · Topics
"Success leaves clues" You gotta start somewhere and you can start at page 1 or page 10 and that time, energy and cost you saved starting 9 pages later can be put into making whatever it is you're building better than the original. The US, and every other country, is full of derivatives or straight up copies. No one is getting super mad at the generic cheerios at the grocery store. It's hypocrisy.
View on HN · Topics
Tell me, where did OpenAI and Anthropic got their training data? From public sources using legitimate means? Don't make me laugh.