Skip to contentSkip to stories

Updated

#Multimodal

Showing low-relevance items too. Hide low-relevance items

Sep 10

Sep 10Thu
  1. DeepSeekAI score72

    DeepSeek V4.1-Flash goes live on its API with native multimodal support

    AIDeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.

  2. DeepSeek API NewsAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

Sep 8

Sep 8Tue
  1. Cohere · new models on Hugging FaceAI score38

    Cohere releases Tiny Aya Base 32K, a 3.35B multilingual model with 32K context

    AICohere Labs has released Tiny Aya Base 32K, an open-weights pretrained model with 3.35 billion parameters and a 32K context window. The model covers 70+ languages, including many lower-resourced ones, and is designed for downstream adaptation and long-context research. It is a base model that has not been instruction-tuned, and it is licensed under CC-BY-NC.

  2. Xiaomi MiMoAI score52

    Xiaomi MiMo Desktop enters invite-only beta as a desktop agent

    AIXiaomi MiMo has launched MiMo Desktop in invite-only beta, a desktop agent that turns Office files, images, video, audio, and zips into finished, editable output. Invitees also get limited access to next-gen MiMo models, and the post lists features including live previews, region-based editing with versioned rollback, automatic model routing, and browser and computer use with record and replay.

  3. NVIDIA · new models on Hugging FaceAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 4

Sep 4Fri
  1. BAAI · new models on Hugging FaceAI score26

    ConsiSpace: BAAI and Peking University release geometry-consistent video spatial reasoning model

    AIBAAI and Peking University researchers released official weights for ConsiSpace, a geometry-consistent multimodal framework for spatial reasoning in long-form visual observations. The model is described in the paper "ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning" (arXiv:2607.17599).

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. Midjourney UpdatesAI score52

    Midjourney's alpha adds v8.2 edit model with lightbox editor

    AIMidjourney's alpha site now runs the new v8.2 edit model, with an editor built into the lightbox. Users can edit images with plain-text instructions, attach up to 4 reference images, and view all session edits in one place. The update also adds an early Change Style feature, and the team says speed and error messaging have improved, while drag and drop and the prompt bar are still in progress.

  2. BAAI · new models on Hugging FaceAI score25

    BAAI Releases Recon2Reason-Reasoning-4B, a Spatial Reasoning Vision-Language Model

    AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.

Sep 2

Sep 2Wed
  1. Understanding AI (Timothy B. Lee)AI score62

    How Google's RT-2 set the template for today's robotics models

    AIGoogle's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

  2. Google AI DevelopersAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post

Sep 1

Sep 1Tue
  1. Google AI StudioAI score75

    Google adds agentic video understanding to Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite

    AIGoogle AI Studio says agentic video understanding is now available across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the Gemini API. The company reports cost reductions of up to 66%, token consumption reductions of up to 88% and accuracy gains of up to 7% on standard video benchmarks. Developers enable it by setting processing to "agentic" in the API configuration, at standard token pricing.

    Why it matters: The source gives concrete cost and token figures and explains how the agentic loop replaces fixed-rate frame ingestion, helping developers weigh it against their current video pipelines.

  2. Google AI StudioAI score62

    Google AI Studio introduces agentic video understanding with Gemini

    AIin which the model decides what to watch, at what speed, and through which modality. It fetches only the moments and signals it needs instead of ingesting media at a fixed frame rate. The post says this cuts costs by up to 66% and token consumption by up to 88% while boosting accuracy, and it is available now via the Gemini API and in AI Studio.

    Video from @GoogleAIStudio's post
  3. Google · new models on Hugging FaceAI score44

    Google Releases GNM v3.0, an Open 3D Parametric Model of the Human Head

    AIGoogle has released GNM v3.0, a parametric 3D statistical model of the human head, with weights published on Hugging Face and Kaggle under the Apache 2.0 license. The model gives controllable identity, expression, head pose, and internal anatomy including eyeballs, teeth, and tongue, and supports NumPy, JAX, PyTorch, and TensorFlow backends.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. Microsoft ResearchAI score45

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Population-Scale Research

    AIMicrosoft Research released GigaPath-Flash and GigaTIME-Flash, efficient pathology foundation models built on a distilled ViT-S backbone and released under the Apache 2.0 license. GigaPath-Flash, with 22M-parameter tile and 21M-parameter slide encoders, reportedly scores within 3% of the original GigaPath on PANDA and EBRAINS benchmarks at roughly 50 times less compute. The models are research tools, not validated for clinical use.

  2. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 27

Aug 27Thu