IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining

| Source: Apple ML Research

Tags: structured pruning, Apple ML Research, model compression, efficient inference, LLM optimization, pretraining

Apple ML Research presents IDEA Prune: an integrated enlarge-and-prune pipeline that outperforms training target-size models from scratch by pretraining a larger model then compressing it via iterative structured pruning under a single cosine annealing schedule — validated at production scale compressing 2.8B to 1.3B parameters with up to 2T pretraining tokens.

Details

IDEA Prune, published in August 2026 by researchers from Apple with affiliations at Georgia Tech and UT Austin, addresses a practical gap in structured model pruning pipelines. Previous enlarge-and-prune approaches treat the three phases — oversized model pretraining, structured pruning, and recovery training — as sequential independent steps with incompatible learning rate schedules. The naive approach causes knowledge loss at recovery time when the learning rate spikes upward again after pruning. The paper's core contribution is integrating all three phases under a single cosine annealing learning rate schedule, eliminating the damaging LR spike. This is paired with a novel iterative structured pruning method that removes parameters gradually across multiple steps rather than in one pass — allowing surviving neurons to progressively redistribute model capacity and enabling smoother compression with better retained performance. The critical empirical question the paper addresses: is it worth pretraining an enlarged model that will never be deployed? The experiments — conducted at 2.8B → 1.3B parameter compression with up to 2 trillion pretraining tokens — show the answer is yes. The enlarged starting point survives aggressive compression better than a model trained at target size from scratch, both in token efficiency and in final pruned model performance. This is directly relevant to practitioners compressing models for inference-constrained environments — edge deployment, cost-optimized serving, or latency-sensitive production applications. The technique is framework-agnostic and applicable to any LLM pretraining pipeline where a larger intermediate checkpoint can be maintained.