Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation
Heyuan Yao, Chutong Gao, Yuan Lyu, Izzy Grosof, David Simchi-Levi
Abstract
The major workloads in modern large language model (LLM) serving systems have shifted from single-shot LLM calls to multi-turn conversations, where new responses are generated based on the whole conversation history across all previous turns. The hit ratio, i.e., the average fraction of KV caches accessed directly from existing caches stored in high-bandwidth memory (HBM), is hence a crucial metric that governs system performance. Estimating the hit ratio is a highly nontrivial task due to the complex system dynamics, where the KV cache prefixes grow with turns and some must be evicted due to finite memory capacity. We formulate the system as a multi-turn conversation model under the least-recently-used (LRU) policy. Through a mean-field asymptotic framework, we prove that as the conversation arrival rate and the memory capacity grow proportionally to infinity, the hit ratio converges to a closed-form limit. Based on the characterization of the limit, we further propose a practical hit ratio estimator, and validate its accuracy by real LLM serving experiments on the Qwen3-8B model implemented on Ascend NPUs. Our results provide a theoretical foundation for the analysis of multi-turn LLM serving systems and a practical guideline for memory capacity provisioning.
Create a lesson
Related papers
The Price of Remembering: A Calibrated Energy Law for Computation
Mohamed Amine Bergach
DART: Aiming for Tail-Delay Control in Reconfigurable Networks
Hossein Mohammadalizadeh, Holger Karl
Spectral Analysis for Sparse Matrix Computation: Insights and Potential
Ruifeng Zhang, Xipeng Shen
Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton et al.
FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval
Long Yang, Yu Mao, Yuchen Shao et al.
Adaptation Fidelity of SPEC CPU2026
Doa'a Al-Otoom, Mahesh Madhav