RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
| Source: arXiv AI
Tags: LLM inference, DeepSeek V4, inference optimization, throughput, KV cache, AI infrastructure, NVIDIA B300
RoofLang is a domain-specific language for AI-driven LLM inference architecture search — revealing that DeepSeek V4 achieves 3.5-39.5x higher peak decode throughput than competing models, and enabling an agent to discover new architectures improving DeepSeek V4 Pro by 6.23-50.1% on NVIDIA B300.
Details
Current AI-driven LLM inference optimization is almost entirely profiling-based: run the existing stack, measure bottlenecks, suggest parameter tweaks. The fundamental limit is that you can only find architectures within reach of your current implementation.\n\nRoofLang breaks this constraint with a domain-specific language (DSL) providing three capabilities: a general workload representation, a verifiable mutation space, and an implementation-independent evaluator. By reasoning about roofline arithmetic rather than profiling output, an optimizer agent can explore architectures that would require software changes to implement.\n\nThe first major finding: analyzing DeepSeek V4-series models through RoofLang reveals they achieve 3.5-39.5x higher peak decode throughput than other representative models — a gap that is disproportionate to parameter count and explained largely by compact KV-cache designs that support larger batches and reduce memory traffic.\n\nThe second finding: a persistent optimizer agent running against RoofLang discovered new architectures that improve both throughput and interactivity of DeepSeek V4 Pro on NVIDIA B300 by 6.23-50.1% over the existing configuration.\n\nFor infrastructure teams: this is a tool for principled architecture search rather than tuning. The RoofLang DSL approach could significantly change how inference systems are designed at scale.