TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 1 October 2026, 01:00 UTC

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

What
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
Who
arxiv.org
When
30 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.36322
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. It comes from a paper posted to arXiv on 30 September 2026. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.

Why it counts

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.