Skip to content
Read the original: Moonshot AI (Kimi) · new models on Hugging Face· Published Pick60/100AI score60/100

Moonshot AI releases Kimi Linear 48B hybrid linear attention models on Hugging Face

Original titlemoonshotai/Kimi-Linear-48B-A3B-Base

AISummary

Moonshot AI released Kimi Linear, a hybrid linear attention architecture with 48B total and 3B activated parameters and a 1M-token context length, on Hugging Face. The model card reports up to 6.3x faster TPOT than MLA at 1M tokens and up to 75% lower KV cache needs, and says it outperforms full attention on long-context and RL-style benchmarks.

AIWhy it matters

The model card gives concrete long-context speed and memory figures for a hybrid attention design, useful for judging whether linear attention can replace full attention in practice.

Read the original huggingface.co

Source: Moonshot AI (Kimi) · new models on Hugging Face · huggingface.coPublished · added here