Beating AI News Flash: ByteDance's Seed team has discovered a regular weakness in DeepSeek-V4's long-text processing. Researchers asked the model to complete a piece of code; the code itself did not change, but simply extending an irrelevant comment before it caused the model's answers to alternate between correct and incorrect. The V4-Flash base version repeats a cycle every 4 tokens, and V4.1-Flash shows a similar phenomenon, with the cycle becoming 2 tokens.
The issue is related to DeepSeek's "chunked KV cache compression," which saves memory. The model merges and stores information from several consecutive tokens to reduce the cost of processing long text. But this compression causes information at different positions to be retained to different degrees, making content at certain positions harder to retrieve accurately later. ByteDance calls this periodic retrieval difference "phase sensitivity."
In a retrieval test of about 128,000 tokens, the V4-Flash base version showed a maximum accuracy difference of 40.2 percentage points across different positions. After post-training, the gap narrowed to 19.1 percentage points; V4.1-Flash further narrowed it to 6.1 percentage points, but periodic fluctuations still remained. ByteDance also trained multiple control models from scratch and found that the fluctuation cycle matched the compression stride, while control models that did not use chunked compression did not show the same pattern.

