Zhang, H., Xia, C. and Wang, Z. orcid.org/0000-0001-6157-0662 (2026) KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference. In: MobiSys '26: Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services. MobiSys '26: 24th Annual International Conference on Mobile Systems, Applications and Services, 21-25 Jun 2026, Cambridge, UK. . Association for Computing Machinery (ACM), New York, NY, United States, pp. 723-736. ISBN: 979-8-4007-2027-7.
Abstract
Language models (LMs) are enabling a new class of mobile and embedded AI applications, including meeting and video summarization, document analysis, and other multi-input workloads that require long-context reasoning. Running these models locally improves privacy, enables offline use, and lowers serving cost. However, on-device long-context inference is fundamentally constrained by a memory capacity wall: the key-value (KV) cache grows linearly with context length and batch size, quickly exceeding available memory. Existing KV-cache offloading schemes are designed to transfer cache data from GPU memory to CPU memory; however, they are not suitable for embedded and mobile systems, where the CPU and GPU (or NPU) typically share unified memory, and secondary storage offers limited, highly asymmetric I/O bandwidth. We present KVSwap, the first KV-cache management framework for resource-constrained mobile systems that uses storage, rather than CPU RAM, as the KV-cache backing store for long-context inference. KVSwap stores the full cache on disk, uses compact in-memory metadata to predict and preload future accesses, overlaps computation with hardware-aware storage transfers, and optimizes read patterns for device characteristics. Our evaluation shows that across representative LMs and storage types, KVSwap delivers higher throughput under tight memory budgets while maintaining generation quality over existing KV cache offloading schemes.
Metadata
| Item Type: | Proceedings Paper |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2026 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution International 4.0 License. |
| Keywords: | large language models; long context inference; mobile system |
| Dates: |
|
| Institution: | The University of Leeds |
| Academic Units: | The University of Leeds > Faculty of Engineering & Physical Sciences (Leeds) > School of Computing (Leeds) |
| Funding Information: | Funder Grant number EPSRC Accounts Payable EP/X037304/1 Royal Society *** USE 813030 *** IF\R1\251008 |
| Date Deposited: | 17 Apr 2026 09:26 |
| Last Modified: | 13 Aug 2026 10:53 |
| Status: | Published |
| Publisher: | Association for Computing Machinery (ACM) |
| Identification Number: | 10.1145/3745756.3809234 |
| Related URLs: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:240121 |

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)