AI & Computingpreprint2026-08-11

Critical-Span Retention and Query-Unknown Discovery for Long-Context KV Compression

Open access0 citations

Abstract

Long-context inference is limited by KV-cache memory. This preprint investigates which token spans must be retained for retrieval-critical tasks and whether online compression fails because of insufficient cache capacity or weak query-unknown discovery. Controlled multi-seed experiments identify a critical-span retention mechanism and isolate an online discovery gap. A lightweight surface-novelty policy closes much of this gap on the controlled retrieval suite through 40k context length while maintaining a measured peak cache of approximately 1k tokens. Experiments on a fixed public LongBench slice expose a less favorable but explicit peak-quality tradeoff. The accompanying open-source implementation, experimental protocols, results, and reproduction scripts are available from the project repository. Code and reproduction materials: https://github.com/nilsperssonsuorra/nearlossless-context

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-11

Authors: Nils Persson Suorra