Data-free On-policy Distillation

| Source: arXiv AI

Tags: knowledge distillation, LLM post-training, on-policy distillation, data efficiency, model training, fine-tuning

On-policy LLM distillation is nearly data-indifferent: just 8 prompts match a 17k-problem dataset, domain doesn't matter, and teachers self-generating their own training questions outperform real external datasets — a finding that rewrites assumptions about post-training data pipelines.

Details

On-policy distillation (OPD) — where a student model learns from a teacher's real-time corrections during training — is standard in frontier post-training pipelines. But nobody had systematically tested how much the training data itself actually matters.\n\nLi et al. run controlled experiments across two teacher-student pairs and find a striking result: OPD is nearly indifferent to its data. Eight prompts match the performance of a 17,000-problem dataset. Three independently-built datasets with substantially different difficulty levels and teacher-student KL divergences produce nearly identical training curves. Swapping math problems for competitive programming recovers over 90% of in-domain performance.\n\nThe explanation: in OPD, each prompt keeps generating new teacher correction as sampling continues. The marginal value of additional prompts collapses after eight. What transfers is the teacher's mode of reasoning, not domain-specific knowledge.\n\nThe paper then demonstrates Data-free OPD (DF-OPD): the teacher writes its own training questions under a simple prompt, with no external data or filtering. DF-OPD matches real data performance — and 1,000 self-generated questions close 98.5% of the available headroom in multi-teacher distillation.