MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
| Source: arXiv AI
Tags: LLM architecture, token embeddings, Llama-3, Qwen3, MobileLLM, sparse lookup
MoME replaces LLM token embeddings with context-aware mixtures of M memory slots, letting the same surface token retrieve different embeddings depending on its meaning — improving over fixed-entry baselines on Llama-3, MobileLLM, and Qwen3 backbones.
Details
Standard token embedding tables assign each token a fixed vector regardless of context — python the programming language and python the snake share one memory row. Mixture of Memory Embeddings (MoME) replaces each token entry with M slots and uses a learned gate over the hidden state to select which slots to read at each position, making token memory context-sensitive. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over three baselines (Value Embedding, Bigram, STEM) in both iso-parameter and iso-training-FLOP settings, and shows a more promising memory-size scaling trend at sub-billion scale. Training and inference efficiency are maintained. Qualitative routing analyses on polysemous tokens suggest the learned mixture achieves a degree of semantic interpretability — the same surface token is dispatched to different slots under different senses. For researchers working on efficient LLM architecture or embedding augmentation, MoME offers a practical and interpretable way to add semantic richness to token lookups without increasing backbone size.