Skip to contentSkip to stories

Updated

#Multimodal

Oct 5

Oct 5Mon
  1. Liquid AIAI score37

    Liquid AI's d1 decision model adds vision, rivaling GPT-6.1 Sol at lower cost

    AILiquid AI released d1 with vision support, accepting images, text, or both as inputs. In tests on six real applications, d1 matched or beat GPT-6.1 Sol on four while costing 19x to 200x less than both GPT-6.1 Sol and Claude Opus 5.5. It returns probabilities for yes/no, choice, or score questions in one forward pass, with text decisions in 200 to 300 ms.

  2. Liquid AIAI score36

    Liquid AI's d1 model inspects parts from camera images with 85-97% accuracy

    AILiquid AI's vision-enabled decision model d1 inspects parts directly from camera images and is described as the best such model currently on the market. It reaches 85% to 97% accuracy across four VisA inspection tasks covering circuit boards, candles, cashews, and chewing gum. It understands each task from a short description without task-specific training.

  3. Liquid AI · new models on Hugging FaceAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.

  4. Liquid AI · new models on Hugging FaceAI score67

    Liquid AI releases d1-3B, a 3B multimodal decision model for edge deployment

    AILiquid AI has released d1-3B, a 3B parameter multimodal model post-trained to return calibrated, typed answers to yes/no, choice, and score questions in one forward pass. The source reports a Decision Index 0.2.1 score of 48.57, the highest among models under 10B in its table, and 8 ms per decision on an NVIDIA RTX 4090.

    Why it matters: The source gives benchmark scores against named peer models and edge latency figures across several hardware targets, helping readers judge fit for on-device decision pipelines.

Oct 4

Oct 4Sun
  1. Liquid AI BlogAI score70

    Liquid AI releases d1 decision model with image input support

    AILiquid AI introduces d1, its first decision model, now accepting both text and images. The company says d1 matches or beats GPT-6.1 Sol on four of six tested applications, at 19x to 200x lower cost and with faster answers on every task. d1 is available on the Liquid AI API and through Vercel and OpenRouter, with text-only support on those two platforms for now.

    Why it matters: The post gives benchmark comparisons against named models along with per-token pricing and latency figures, which makes the cost and speed tradeoff checkable.

  2. Kling AIAI score36

    Kling 4.0 powers "The Beat," a 30-second continuous-shot short film

    AIKling AI used its Kling 4.0 model to produce "The Beat," a short film built around a 30-second continuous shot and surpassing 5 million impressions across social platforms. The model's native 30-second generation, Omni Reference supporting up to 15 multi-modal references, Multi-Keyframe control for up to 10 keyframes, and 10-bit HDR output shaped the film's continuity, consistency, and color. The post walks through these features shot by shot.

Oct 3

Oct 3Sat

Oct 2

Oct 2Fri
  1. Google · AI blogAI score58

    Google recaps September 2026 AI launches, led by Gemini 4 Argon

    AIGoogle's September 2026 roundup highlights Gemini 4 Argon, a frontier model with a 1-million-token output limit aimed at complex tasks such as cybersecurity defense. Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with developer, enterprise, and consumer access to follow after guardrail feedback. The post also covers Gemini 3.8 Flash, Connected Apps in Gemini, and WeatherNext 3.

  2. Kling AIAI score52

    Kling 4.0 Enters Closed Beta With Stable Motion and 30-Second Takes

    AIKling 4.0 is in closed beta, with an official launch planned for October, and supports video up to 30 seconds long with up to 4K resolution and 10-bit HDR output. Creator Johnson Sheng reports stable dynamic motion, consistent characters and props across shots, and an unedited 30-second fight sequence, while noting that Kling 4.0 is aimed at commercial production.

Oct 1

Oct 1Thu
  1. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.

  2. NVIDIA · new models on Hugging FaceAI score44

    NVIDIA releases PixelUMM, an encoder-free model for pixel-space image and video tasks

    AINVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space. It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.

  3. Google · Gemini appAI score60

    Google launches Guided Vision in Gemini Live for blind and low-vision users

    AIGoogle is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.

    Why it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.

  4. Meta NewsroomAI score22

    Ranveer Singh Becomes Ray-Ban and Ray-Ban Meta Brand Ambassador in India

    AIMeta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.

  5. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick Says Agent Coordination Is Easier Than Expected

    AIEthan Mollick says he was wrong to think coordinating AI agents would require careful human-designed management structures. He points to personal agents like dots and Muse, and to a swarm of thousands of OpenAI agents that solved a Navier-Stokes problem in 88 hours with thin coordination. He argues many management problems stem from human limits, which agents lack, so people should mainly guide direction while agents handle organizing.

  6. Manus BlogAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.