Skip to content
Read the original: DeepSeek· Published 38/100AI score38/100

DeepSeek unveils 552B MoE model with asymmetric encoder-decoder design

Original title🧠 Asymmetric architecture. More intelligence, less cost.

AISummary

DeepSeek has introduced a 552B-parameter MoE model built on a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The company says new pre-training methods and larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

Read the original x.com

Source: DeepSeek · x.comPublished · added here