LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
| Source: arXiv AI
Tags: LLM inference, layer skipping, LoRA, Qwen2.5, efficient inference, transformer
LayerRoute skips the same 9 middle transformer layers (8-16) in Qwen2.5-0.5B across all 10 training seeds, delivering 1.02-1.06x wallclock speedup with improved perplexity — but results are on a 0.5B model only and speedup is modest.
Details
Adaptive layer-skipping is an appealing approach for inference efficiency: skip redundant transformer layers for simple inputs while using the full model for complex ones. LayerRoute implements this via a lightweight per-layer binary router (~21.5K parameters each) trained via straight-through estimator, combined with LoRA adapters (rank 8, ~1.08M total) that compensate for skipped computation. The most interesting finding is the extreme consistency: across 10 independent training runs with different random seeds, the router converged to exactly the same skip pattern — layers 8-16 are skip-eligible in every run. This suggests the redundancy structure in this specific model is robust and deterministic, not an artifact of training stochasticity. Measured speedup is 1.02-1.06x (mean 1.04x), which is modest. The authors verify this is wallclock speedup, not theoretical FLOP reduction. Perplexity improves across all 10 seeds, suggesting LoRA adapters more than compensate for skipped layers. Training takes under 7 minutes on a single A100. Importantly, all results are on Qwen2.5-0.5B-Instruct. Whether the consistent skip-pattern finding holds for larger models, or whether speedup scales favorably, is an open question. The 1.04x mean speedup is unlikely to be production-relevant on its own.