❯ ByteDance’s Seed team finds chunked KV cache compression causes “phase sensitivity”, with long-context retrieval accuracy varying by up to 40.2 points
Move the information, change the accuracyByteDance’s Seed team submitted a paper to the preprint server arXiv in late September studying a side effect of a long-context optimization technique. According to PEdaily, in a 128K long-context retrieval test the researchers found that simply moving the same piece of information changed retrieval accuracy in DeepSeek-V4 series models by up to 40.2 percentage points, rising and falling on a 4-token cycle.
Position inside the compression window mattersThe cause points to chunked KV cache compression. As ITHome explains, models must store large amounts of history when processing long text, and this technique compresses consecutive tokens at a fixed stride into fewer cache entries to save memory and computation. The side effect is that each token gains a new attribute: its position within the compression window. The team calls the fact that the same information is easier or harder to retrieve at different positions “phase sensitivity.”
Newer versions have narrowed the gapThe paper evaluated several versions. The gap was largest in the base version of DeepSeek-V4-Flash; later versions narrowed it to 19.1 points for V4-Flash-0731 and 14.8 points for V4-Pro-0813, and the newer V4.1-Flash-0910 brought it down to 6.1 points, with the cycle shifting from 4 tokens to 2. The team also built several compression schemes on the Qwen3-0.6B architecture as controls and saw the same effect, indicating it stems from the chunked-compression design itself.
Averages hide itStandard evaluations pool many test results into one average, so periodic highs and lows cancel out. The paper recommends that models using this kind of compression be tested with information placed at different phases. For developers, it offers a testable explanation for why the same question over a long document sometimes works and sometimes does not.
▮ SIGNALAn engineering optimization made to save memory can leave regular weak spots in a model, and comparing long-context ability calls for measurements finer than an average score.