ByteDance Researchers Discover Phase Sensitivity Issue in Chunked KV-Cache Compression
Read more
Pandaily
pandaily.com

ByteDance Researchers Discover Phase Sensitivity Issue in Chunked KV-Cache Compression

Researchers from ByteDance Seed published a scientific paper in which they identified a periodic vulnerability in language models that use fixed-block compression of key-value caches (KV-cache). This approach is used, for example, in the DeepSeek-V4 model to reduce memory and attention costs when handling long contexts.

In the article titled 'Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression,' posted on arXiv on September 28, the team reported that the same information can be easily extracted at one position but is difficult to access at another. This gap may go unnoticed when using standard evaluation metrics.

The chunked compression method groups sequential tokens into windows and compresses each window into a smaller number of cache entries with a fixed stride. This assigns a new coordinate to each token, which the authors call its phase—that is, its position relative to the window boundaries. Even a small shift of the input data by a few tokens can change the group of tokens that are compressed together without changing their actual content.

The team tested basic and fine-tuned versions of DeepSeek-V4-Flash and DeepSeek-V4-Pro, as well as fine-tuned DeepSeek-V4.1-Flash. The testing was conducted on a 'needle in a haystack' search task with a volume of 128 thousand tokens, while keeping the query length and absolute position constant. Accuracy varied by up to 40.2 percentage points depending on the target element's phase in the basic checkpoints. Although fine-tuning improved accuracy and narrowed the gaps, they remained significant for the V4 models.

The DeepSeek-V4.1-Flash model, which uses a stride of two instead of four, demonstrated the smallest gap, although the periodic pattern persisted. To find the cause, researchers pre-trained a family of small transformers from scratch based on the Qwen3-0.6B backbone, changing only the KV compression design. Each compressed variation showed periodicity corresponding to its stride, including a variant that simply averages tokens in each window, whereas the full base attention models did not exhibit this.

Interventions in attention heads suggest what the authors termed 'phase specialization,' as different components process information at different phases. They note that these mechanistic findings are strongest in controlled settings and do not establish a universal cause for large models.

The practical significance of this work relates more to evaluation methods than to any single model. The paper argues that models using KV chunked compression should be evaluated considering phases. For instance, this can be done by shifting the input data by a few tokens while keeping the requested information unchanged, as high average accuracy can coexist with systematic positional failures.

Eight authors listed ByteDance Seed as an affiliation. Three also listed Princeton University, Stanford University, or UC Berkeley, and the paper states that the main contributors performed this work during internships at ByteDance Seed. Corresponding authors are Yiyuan Ma and Xin Dong.

Popular