Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. Meta NewsroomAI score22

    Ranveer Singh Becomes Ray-Ban and Ray-Ban Meta Brand Ambassador in India

    AIMeta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.

  2. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick Says Agent Coordination Is Easier Than Expected

    AIEthan Mollick says he was wrong to think coordinating AI agents would require careful human-designed management structures. He points to personal agents like dots and Muse, and to a swarm of thousands of OpenAI agents that solved a Navier-Stokes problem in 88 hours with thin coordination. He argues many management problems stem from human limits, which agents lack, so people should mainly guide direction while agents handle organizing.

  3. Manus BlogAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

Sep 30

Sep 30Wed
  1. Google FlowAI score38

    Google's Gemini Omni Flash guide offers prompting tips for Flow videos.

    AIGoogle Flow publishes a guide to creative prompting with Gemini Omni Flash, covering video generation for films, marketing, and visual assets. The guide recommends high-level constraints, first and last frame visual anchors, tagged image, video, and storyboard ingredients, and granular mid-scene pacing edits. It also suggests transferring style and motion from reference images and videos.

  2. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  3. IdeogramAI score23

    Ideogram 4.5 Performs Strongly Across General Image Editing Tasks

    AIIdeogram 4.5 was built for precise, targeted editing but also performs very well across general editing tasks. Design Arena ranks it 15th in Image Editing with an Elo of 1250, placing it in the same performance band as MAI-Image-2.6 and Gemini 3 Pro Image Preview. It is especially strong at typography edits, such as modifying text in infographics.

  4. IdeogramAI score38

    Ideogram 4.5 launches as a precise image edit model

    AIIdeogram released Ideogram 4.5, which it calls the most precise edit model, claiming it avoids the artifacts, pixel shifts, and color changes that leading models add with each edit. The company says this eliminates artifact buildup and makes multi-turn editing possible. It is live in Ideogram, via the API, and with launch partners, with open weights promised soon.

    Video from @ideogram_ai's post
  5. SenseTimeAI score23

    SenseTime previews Dynamic Design, animating static images with SenseNova 6.8 Flash

    AISenseTime previewed Dynamic Design, powered by SenseNova 6.8 Flash, which turns static images into animated visuals. The system decides which elements stay static and which to animate, chooses HTML/CSS, SVG, transparent images, or video for each element, and choreographs text reveals and subject motion. SenseNova 6.8 Flash is coming soon, and SenseTime is offering a limited beta.

    Video from @SenseTime_AI's post
  6. Kling AIAI score35

    Kling 4.0 full-powered version showcased in a short film demo

    AIKling AI showcased a short film generated by the full-powered KLING 4.0, which the source describes as the full-powered version. The background post from @hq4ai says the film was made with all-round reference generation, runs a native 30 seconds in 21:9 cinematic format, and offers clearer visuals and sound with more precise lip-sync than Flash. The same post states Kling 4.0 will launch in October.

  7. ModelScopeAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post
  8. Kling AI BlogAI score49

    Kling 4.0 Extends Native Video to 30 Seconds With Up to 10 Keyframes

    AIKling 4.0 extends native single-pass video generation from 15 to 30 seconds and adds Multiple Keyframes supporting up to 10 keyframe images, versus Start & End Frames in Kling 3.0. It also expands reference inputs to up to 15 combined assets, including up to 5 videos totaling 30 seconds, and adds 10-bit HDR at 1080p and 4K. The all-new Kling 4.0 will officially launch in October, and Kling 4.0 Flash became available to a limited group of early-access users on September 28.

Sep 29

Sep 29Tue
  1. SGLangAI score36

    SGLang adds native decision API for classification and scoring models

    AISGLang says it turned Qwen3.8-27B into a multimodal decision model that beat Pokémon FireRed's Elite Four and champion with sub-100 ms decisions from live game state. It introduces a native /v1/decisions endpoint for turning LLMs and VLMs into classification and scoring models. A /v1/systemone endpoint is also added so Jev-like open models can work with the TypeSafe SDK.

    Video from @sgl_project's post
  2. Jerry LiuAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post
  3. Google ResearchAI score35

    Google Research unveils Diffusion Controller for steering AI image generation

    AIGoogle Research introduced Diffusion Controller, a framework that treats image generation as a continuous control problem rather than separate inference-time guidance and fine-tuning fixes. Its lightweight add-on "steering damper" network keeps the base model frozen and works on black-box or gray-box models, and it outperformed the industry standard on human preference matching. In a Stable Diffusion v1.4 test, the fully unlocked version achieved a 90% win rate over the baseline.

  4. Microsoft ResearchAI score34

    Microsoft Research unveils Quine, an early multimodal world model of biology

    AIMicrosoft Research has introduced Quine, an early-stage research effort to build a multimodal world model of biology that connects insights across biological scales and modalities. The system is designed to help scientists computationally search a space far larger than intuition allows and prioritize hypotheses before lab testing. Experimental results are meant to feed back into the model and sharpen future research directions.

    Video from @MSFTResearch's post
  5. ModelScopeAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

    Image from @ModelScope2022's post
  6. OpenBMBAI score34

    MiniCPM-o 4.5 now runs in SGLang Omni v0.1.7 for developers

    AIOpenBMB announced that MiniCPM-o 4.5 is now supported in SGLang Omni v0.1.7, giving developers more flexibility to run and build with the model. The background release notes add that MiniCPM-o 4.5 brings multimodal input and speech output to the runtime. MiniCPM-o and MiniMax-Music3 also gained Intel XPU support in the same release.

  7. Luma AI NewsAI score22

    AI Photo Editing Prompt Formula Preserves Color, Light, and Skin in Campaign Edits

    AIThe article presents a four-part prompt structure (action verb, target element, desired result, protection instructions) for AI photo editing, saying it preserves approved work across platforms. It identifies three common failure causes: unmatched light direction, stacked edits in one prompt, and vague visual language. It states that simple skin retouching takes 2-3 minutes versus 15-30 minutes manually.

Sep 28

Sep 28Mon
  1. LlamaIndex 🦙AI score30

    LlamaIndex says frontier VLMs still struggle parsing tax and W-series forms

    AILlamaIndex argues that frontier vision-language models still fail on real forms such as W-2s, 1040s, W-9s, and scanned W-4s, because forms require detecting every field, preserving section hierarchy, linking values to their exact boxes, and reading handwriting and checkmarks. The company's blog post details these failure modes and presents a custom cookbook for LlamaParse as a cheaper way to handle such forms.

    Image from @llama_index's post
  2. Google · Gemini appAI score38

    See what 4 builders are making with Gemini 3.8 Flash

    AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.

  3. Kling AIAI score42

    Kling 4.0 Flash launches now for Ultra Yearly subscribers; Kling 4.0 arrives October

    AIKling AI says its Kling 4.0 Flash is live now for Ultra Yearly subscribers, with the full Kling 4.0 coming this October. The update advertises up to 4K resolution, 10-bit HDR output, stereo audio, and native 30-second generation. It also adds Omni Reference supporting up to 15 multimodal references and multi-keyframe control with up to 10 keyframes.

    Video from @Kling_ai's post
  4. TechNode · AIAI score60

    Sanxingdui: Future Past, China's AI-produced theatrical film, releases October 23

    AIBona Film Group announced that Sanxingdui: Future Past, a 100-minute film using AI throughout production, will screen nationwide on October 23, 2026. The production team says AI handled tasks like image generation while over 100 professionals kept creative control, and it took two years and over 1.2 million source images to maintain consistency across the film.

  5. ModelScopeAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post

Sep 27

Sep 27Sun
  1. DeedyAI score34

    Deedy urges explainer videos for every open source repo, citing SQLite example

    AIDeedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.

    Video from @deedydas's post
  2. Exponential ViewAI score44

    DeepMind Essay Argues AGI Will Emerge Through Collective Cooperation Among AI Agents

    AIDeepMind has published an essay arguing that AGI will emerge through "cooperative interactions among models, tools, institutions, and human participants" rather than from a single winning AI. The commentary supports the collective framing but rejects treating AI agents as having their own theory of mind, arguing that creating new moral subjects should remain humanity's remit.

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

Sep 25

Sep 25Fri
  1. Google AIAI score57

    Google AI lists weekly releases including Gemini 3.8 TTS, Live Avatar, and Project Suncatcher

    AIGoogle AI's weekly roundup lists Gemini 3.8 Flash TTS and Flash-Lite TTS as expressive audio generation models. It also announces Gemini 3.8 Live with Live Avatar for near real-time visual conversation and a Live Chat voice feature on the Gemini Notebook mobile app across about 100 languages. Project Suncatcher will launch a prototype satellite to test Google TPUs in orbit and explore solar-powered AI compute in space.