How llm-d makes the most of the hardware you already have

| Source: IBM Research

Tags: llm-d, IBM Research, Red Hat, inference, H100, open-source, agentic AI, GLM-5.2

IBM Research, Red Hat, and Google's open-source llm-d framework ran GLM-5.2 (753B-parameter MoE) on 544 H100 GPUs, serving 3,000 concurrent coding agents at 6.6M output tokens/min — self-hosting costs 5-10x less per token than commercial APIs.

Details

IBM Research, Red Hat, and Google demonstrated that self-hosted open-weight models can match commercial API performance for agentic workloads using their open-source llm-d inference framework. The team deployed GLM-5.2, a 753B-parameter mixture-of-experts model (~39B active parameters), across 544 NVIDIA H100 GPUs — the previous-generation accelerator most enterprise GPU fleets already operate. On representative agentic benchmarks, the deployment delivered over 6.6 million output tokens per minute at peak, serving up to 3,000 concurrent coding agents with zero preemptions. llm-d is architected specifically for agentic traffic patterns: long-context reuse, parallel sub-agent spawning, and bursty generation activity that differs fundamentally from chatbot workloads where generation dominates over context processing. The cost advantage is concrete: at current cloud GPU rental rates, running GLM-5.2 with llm-d costs 5-10x less per token than equivalent commercial API pricing. The savings are largest for input-heavy workloads — exactly the traffic pattern that dominates coding-agent sessions, where agents spend most compute reading and reprocessing existing context rather than generating new text. For enterprises seeking to reduce AI inference costs while keeping proprietary data off external APIs, llm-d provides a benchmarked, production-tested path on hardware they likely already own. The project is fully open-source and jointly maintained by IBM Research, Red Hat, and Google.