Skip to content
Read the original: Ahead of AI (Sebastian Raschka)· Published 62/100AI score62/100

Recent LLM architecture changes that cut long-context KV cache and attention cost

Original titleRecent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

AISummary

Sebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs.

He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4.

The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.

Read the original magazine.sebastianraschka.com

Source: Ahead of AI (Sebastian Raschka) · magazine.sebastianraschka.comPublished · added here