Researchers from ByteDance’s Seed team have identified a mechanism behind inconsistent long-context retrieval in DeepSeek models. Their study found that chunked KV-cache compression makes retrieval accuracy sensitive to where information falls within a compression window, producing differences of up to 40 percentage points across positions.
The technique reduces memory and attention costs by compressing groups of consecutive tokens into fewer cache entries. However, the researchers found that the same information can be easy to retrieve at one position and difficult at another, a pattern they call “phase sensitivity.” They reproduced the behavior in models trained from scratch and found that different attention components specialize in retrieving information from different positions. The findings explain how strong average benchmark results can conceal recurring retrieval weaknesses. [ByteDance Seed research paper]
