Skip to contentSkip to stories

Updated

#Multimodal

Showing low-relevance items too. Hide low-relevance items

Oct 2

Oct 2Fri
  1. Kling AIAI score52

    Kling 4.0 Enters Closed Beta With Stable Motion and 30-Second Takes

    AIKling 4.0 is in closed beta, with an official launch planned for October, and supports video up to 30 seconds long with up to 4K resolution and 10-bit HDR output. Creator Johnson Sheng reports stable dynamic motion, consistent characters and props across shots, and an unedited 30-second fight sequence, while noting that Kling 4.0 is aimed at commercial production.

Oct 1

Oct 1Thu
  1. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.

  2. NVIDIA · new models on Hugging FaceAI score44

    NVIDIA releases PixelUMM, an encoder-free model for pixel-space image and video tasks

    AINVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space. It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.

  3. Google · Gemini appAI score60

    Google launches Guided Vision in Gemini Live for blind and low-vision users

    AIGoogle is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.

    Why it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.

  4. Meta NewsroomAI score22

    Ranveer Singh Becomes Ray-Ban and Ray-Ban Meta Brand Ambassador in India

    AIMeta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.

  5. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick Says Agent Coordination Is Easier Than Expected

    AIEthan Mollick says he was wrong to think coordinating AI agents would require careful human-designed management structures. He points to personal agents like dots and Muse, and to a swarm of thousands of OpenAI agents that solved a Navier-Stokes problem in 88 hours with thin coordination. He argues many management problems stem from human limits, which agents lack, so people should mainly guide direction while agents handle organizing.

  6. Manus BlogAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

Sep 30

Sep 30Wed
  1. Google FlowAI score38

    Google's Gemini Omni Flash guide offers prompting tips for Flow videos.

    AIGoogle Flow publishes a guide to creative prompting with Gemini Omni Flash, covering video generation for films, marketing, and visual assets. The guide recommends high-level constraints, first and last frame visual anchors, tagged image, video, and storyboard ingredients, and granular mid-scene pacing edits. It also suggests transferring style and motion from reference images and videos.

  2. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  3. LlamaIndex 🦙AI score14

    LlamaIndex hosts document-processing events for AI agents in New York and San Francisco

    AILlamaIndex held Tuesday-night events in New York and San Francisco on document processing for AI agents, with the New York room filling a waitlist and San Francisco drawing almost 600 attendees. The talks focused on the problem that agents often receive document text without its layout, so they must guess which figures, such as a monthly rate versus a total on an invoice, mean what.

    Image from @llama_index's post
  4. IdeogramAI score23

    Ideogram 4.5 Performs Strongly Across General Image Editing Tasks

    AIIdeogram 4.5 was built for precise, targeted editing but also performs very well across general editing tasks. Design Arena ranks it 15th in Image Editing with an Elo of 1250, placing it in the same performance band as MAI-Image-2.6 and Gemini 3 Pro Image Preview. It is especially strong at typography edits, such as modifying text in infographics.

  5. IdeogramAI score38

    Ideogram 4.5 launches as a precise image edit model

    AIIdeogram released Ideogram 4.5, which it calls the most precise edit model, claiming it avoids the artifacts, pixel shifts, and color changes that leading models add with each edit. The company says this eliminates artifact buildup and makes multi-turn editing possible. It is live in Ideogram, via the API, and with launch partners, with open weights promised soon.

    Video from @ideogram_ai's post
  6. SenseTimeAI score23

    SenseTime previews Dynamic Design, animating static images with SenseNova 6.8 Flash

    AISenseTime previewed Dynamic Design, powered by SenseNova 6.8 Flash, which turns static images into animated visuals. The system decides which elements stay static and which to animate, chooses HTML/CSS, SVG, transparent images, or video for each element, and choreographs text reveals and subject motion. SenseNova 6.8 Flash is coming soon, and SenseTime is offering a limited beta.

    Video from @SenseTime_AI's post
  7. Kling AIAI score35

    Kling 4.0 full-powered version showcased in a short film demo

    AIKling AI showcased a short film generated by the full-powered KLING 4.0, which the source describes as the full-powered version. The background post from @hq4ai says the film was made with all-round reference generation, runs a native 30 seconds in 21:9 cinematic format, and offers clearer visuals and sound with more precise lip-sync than Flash. The same post states Kling 4.0 will launch in October.

  8. ModelScopeAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post
  9. Kling AI BlogAI score49

    Kling 4.0 Extends Native Video to 30 Seconds With Up to 10 Keyframes

    AIKling 4.0 extends native single-pass video generation from 15 to 30 seconds and adds Multiple Keyframes supporting up to 10 keyframe images, versus Start & End Frames in Kling 3.0. It also expands reference inputs to up to 15 combined assets, including up to 5 videos totaling 30 seconds, and adds 10-bit HDR at 1080p and 4K. The all-new Kling 4.0 will officially launch in October, and Kling 4.0 Flash became available to a limited group of early-access users on September 28.

Sep 29

Sep 29Tue
  1. SGLangAI score36

    SGLang adds native decision API for classification and scoring models

    AISGLang says it turned Qwen3.8-27B into a multimodal decision model that beat Pokémon FireRed's Elite Four and champion with sub-100 ms decisions from live game state. It introduces a native /v1/decisions endpoint for turning LLMs and VLMs into classification and scoring models. A /v1/systemone endpoint is also added so Jev-like open models can work with the TypeSafe SDK.

    Video from @sgl_project's post
  2. Jerry LiuAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post