Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
| Source: arXiv AI
Tags: byte models, tokenization, distillation, small models, scaling laws, Llama
At scale, distilled byte-level models surpass token-based 1B models by up to 8.1% on downstream benchmarks and match token model performance using just one-sixth the training data — challenging the assumption that byte tokenization is inherently less efficient.
Details
Byte-level language models operate over raw bytes (256-token vocabulary) instead of subword tokens (~100K vocabulary), eliminating the need for a tokenizer and simplifying cross-language and cross-modal generalization. The tradeoff has always been efficiency: bytes require more steps per character than tokens. This paper from Marathe et al. (Meta/UW team) runs the first large-scale comparison of distilled token vs. byte models at 1 billion parameters, sweeping up to 1 trillion bytes of training data. The key findings: token models are more efficient in the low-FLOP regime, but byte models eventually surpass them at higher compute budgets, reaching a higher asymptotic performance ceiling. Distilled End-Of-Token-1B models are predicted to outperform distilled Token-1B by up to 4% asymptotically and beat Llama 3.2-1B, Gemma-3-1B, and Gemma 2B by up to 6.5%, 8.1%, and 2.1% respectively. Byte models are also dramatically more data-efficient: they match token model performance at 1/6 the training data. Two logit conversion methods are introduced to enable distillation from token models to byte models. The practical implication: if training budgets are large enough, byte-level models may become the preferred architecture for 1B-class deployments.