DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
| Source: MarkTechPost
Tags: DeepSeek, DeepSeek-V4.1-Flash, MoE, KV Cache, FP4, long-context, inference efficiency, open-source
DeepSeek-V4.1-Flash is a 552B MoE model with 1M-token context window that compresses global KV cache to 890 bytes per token — 437x smaller than DeepSeek-V1 — using FP4 quantization and cross-layer attention reuse. MIT-licensed with vLLM and SGLang support, it directly targets the memory bottleneck in long-context agent deployments.
Details
DeepSeek-V4.1-Flash targets a real production pain point: long-context requests flooding GPU memory with massive KV caches. The model splits its 40-layer backbone into a 20-layer causal encoder and a 20-layer decoder, drawing from the YOCO architecture. Prompt tokens only traverse the encoder; the decoder derives its global KV from the final encoder hidden state rather than recomputing it, cutting prefill compute roughly in half. The cache compression story is equally aggressive. Compressed Sparse Attention 2 (CSA2) assigns each layer one of three roles: Full (computes fresh KV and selects Top-512 indices), Reindex (reuses prior KV but rescores with its own indexer Q), or Reuse (skips scoring entirely, sharing indices from the nearest Full layer). Most layers operate in Reuse mode. The main KV is then quantized to FP4 (E2M1 format with E4M3 scale per 16 channels), halving storage versus V4-Flash FP8. Combined result: 890 bytes per token, 437x smaller than DeepSeek-V1. The model adds 196B Engram parameters alongside the 552B backbone — described as a persistent memory component, though the article is light on details. Open weights ship MIT-licensed on Hugging Face with vLLM, SGLang, and Transformers compatibility. A public API offers three reasoning tiers (low, high, max), enabling immediate use before self-hosting is viable. For teams running long-context agents or RAG pipelines at scale, the KV cache reduction translates directly to infrastructure savings: more concurrent sessions per GPU and cheaper SSD offloading.