Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face
AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.
Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.





