Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache

| Source: arXiv AI

Tags: vLLM, LMCache, GLM-5.3-Flash, KV cache, inference optimization

A targeted fix for GLM-5.3-Flash hybrid-state cache recovery in vLLM+LMCache corrects a token-count mismatch during complete-hit recovery, improving generation agreement from 34/36 to 36/36, while CPU reload cuts time-to-first-token by 46–64% versus cold recomputation across 120 requests.

Details

When vLLM uses LMCache for external KV cache transfers, a subtle bug in GLM-5.3-Flash's hybrid-state model caused cache recovery to succeed at the transfer level while resuming from an inconsistent state: the complete-hit recovery path restored state for the full prompt but the scheduler credited one fewer token.\n\nThis paper documents the diagnostic process and fix: strict-prefix lookup alignment, shared computation corrections, matched checkpoint scheduling, and fixed per-rank kernel configurations for the 45-layer GLM-5.3-Flash model under four-way tensor parallelism (RedHatAI/GLM-5.3-Flash-NVFP4 checkpoint).\n\nIn a nine-length serial workload, agreement with recomputation control improved from 34/36 to 36/36 generations across 64 token IDs. A subsequent performance study across 120 requests showed CPU reload reducing time-to-first-token by 46–64% and total request time by 1.9–7.0% relative to cold recomputation.\n\nThe contribution is narrow: an experimentally validated integration repair for one specific model revision and configuration. The paper explicitly notes it does not establish general determinism or concurrent-serving gains. Useful for teams deploying GLM-5.3-Flash with vLLM caching.