SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation
| Source: arXiv AI
Tags: Qwen3, self-distillation, knowledge distillation, OPSD, fine-tuning, LLM training
SCOPE-OPSD improves on-policy self-distillation by adding a second supervision channel: projecting teacher-student layer discrepancies onto a Fisher-conditioned rank-64 subspace. Tested on Qwen3-1.7B, 4B, and 8B, it beats baseline OPSD in 11 of 12 model-checkpoint combinations without adding rollouts or inference modules.
Details
On-policy self-distillation (OPSD) trains LLMs by scoring student-generated completions with a solution-conditioned self-teacher, but transfers supervision only through next-token probabilities. SCOPE-OPSD explores whether the final-layer discrepancy between teacher and student—ignored by vanilla OPSD—provides a useful additional signal.\n\nThe method projects the teacher-student residual onto a frozen rank-64 factor computed from residual covariance and Fisher sensitivity of the language model head. Critically, it reuses the forward passes already required by OPSD, adding no extra rollouts or inference modules. A matched Random control isolates whether the structured, data-dependent orientation (versus random rotation) drives the gains.\n\nAcross 25/50/75/100 training steps for Qwen3-1.7B, 4B, and 8B, the Structured variant never falls below pure OPSD and exceeds the matched Random baseline in 10 of 12 combinations—with a 1.39 Macro Avg@12 point advantage at step 75 on 1.7B, reproduced across two independent reruns. Results come from Alibaba Cloud and Chongqing Ant Consumer Finance teams.