Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters
Original titlezai-org/GLM-5.3-Flash
AISummary
Z.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters.
The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.
AIWhy it matters
The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.
Source: Z.ai (GLM) · new models on Hugging Face · huggingface.coPublished · added here