PROOF-Gen: From Optimized Data to Better Distillation
| Source: Apple ML Research
Tags: PROOF-Gen, Apple, distillation, tool-calling, Qwen3, Gemma, EMNLP, SFT
Apple researchers' PROOF-Gen recovers 93% of failed teacher trajectories during model distillation by using a per-scenario reflector to generate corrective guidance, boosting Qwen3-4B Pass^1 on tool-calling from 0.132 to 0.529 — a 4x gain — with positive transfer to deployed on-device models across all locales.
Details
Distilling tool-calling capabilities from frontier models into smaller, deployable ones relies on supervised fine-tuning on teacher-generated trajectories. The problem: on τ2-bench, 57% of teacher trials fail, and two-thirds of those are near-misses — trajectories where a single wrong tool call invalidates otherwise correct reasoning. Standard generate-and-filter pipelines discard these failures entirely, wasting compute and leaving hard scenarios perpetually unsolved. PROOF-Gen (Per-scenario Reflective Optimization to Overcome Failed Generation), published by Apple ML Research at EMNLP 2026, addresses this directly. For each failed task, a reflector module analyzes the full execution trace and evaluation feedback, then writes corrective guidance that steers the teacher toward a passing trajectory. Before training, that guidance is stripped — so the student model trains on clean demonstrations with no task-specific scaffold embedded. Results are concrete: 93% of previously failed scenarios recovered. Qwen3-4B-Instruct-2507 trained on augmented data improves Pass^1 from 0.132 to 0.529; Gemma 4 E4B-it gains 7.2pp on BFCL v4 multi-turn. In Apple's production pipeline the method adds +6.3pp goal completion. On-device models see +1.5pp goal completion and +1.7 to +5.0pp on response-quality metrics, with positive transfer in every locale tested (non-English average +1.48pp). For teams running daily or weekly distillation cycles — common in production tool-calling pipelines — PROOF-Gen extracts more training signal from the same frontier-teacher compute by salvaging near-miss failures rather than discarding them.