ByteDance Seed finds phase-dependent retrieval gaps in chunked KV-cache compression
Overview
ByteDance Seed researchers report that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position, a gap their paper says average benchmark scores can hide.
In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by position phase, according to the Pandaily report. The authors say post-training narrowed these gaps but did not eliminate them, and they urge evaluating such models across positional phases rather than relying on aggregate accuracy.
Written by AI from the articles below · updated Oct 8, 10:57 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
- PandailyByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression
AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.
Heat trend
Not enough continuous observations to show a trend yet.