Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. Vercel DevelopersOfficialAI score22

    Liquid AI's d1 model now available on Vercel AI Gateway

    AIVercel says Liquid AI's d1 model is live on AI Gateway under the identifier liquid/d1. The model supports vision inputs for classifying, routing, and scoring decisions.

  2. Artificial AnalysisOfficialAI score32

    HiDream-O1-Video-1.0 ranks #6 on Artificial Analysis image-to-video leaderboard

    AIHiDream-O1-Video-1.0 ranks #6 in Artificial Analysis's Image to Video with Audio leaderboard, just behind Dreamina Seedance 2.0 720p. HiDream says the model generates 1080p videos of 5 to 20 seconds with synchronized audio, priced at $5.80 per minute ($0.10 per second) on the HiHarness API. It is also available in vivago R1 Studio.

    GIF from @ArtificialAnlys's post
  3. TinkerOfficialAI score40

    Tinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash models

    AITinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash, both of which natively accept image inputs and use efficient attention architecture. GLM-5.3-Flash costs 4-5 times less on Tinker than GLM-5.3. Long-context options for Qwen3.5-4B and Qwen3.6-35B-A3B are also live.

  4. StepFunOfficialAI score34

    StepFun's Step 5 Preview free on Nous Portal this week

    AIStepFun's Step 5 Preview, a 600B-total, 27B-active MoE model with 1M context and vision, is free to try on Nous Portal for one week. Nous Research says it scored 33.89 on the Hermes Index, the same score as GPT-6 Luna.

  5. Julien ChaumondXAI score23

    Cloudflare releases clef-omni model on Hugging Face

    AICloudflare has published a new model called clef-omni on Hugging Face, according to a post from Julien Chaumond, who owns the account. The post links to the model page but gives no further details about its size, capabilities, or benchmarks.

  6. ChatGPTOfficialAI score22

    ChatGPT app lets users create dots from phones

    AIOpenAI says users can now create their dot directly from their phone in the ChatGPT app on iOS and Android. The post presents this as the first of several fresh updates for dots, and does not give further details.

    Video from @ChatGPT's post
  7. Nous ResearchOfficialAI score49

    StepFun's Step 5 Preview is free on Nous Portal for a week

    AINous Research says StepFun's Step 5 Preview is free on Nous Portal for the next week. The model is a 600B total, 27B active MoE with a 1M context window and vision support. It scored 33.89 on the Hermes Index, the same score as GPT-6 Luna.

    Video from @NousResearch's post
  8. Cloudflare Blog · AIOfficialAI score55

    Cloudflare releases Clef-omni with audio and video input and cuts Clef-flash price

    AICloudflare releases Clef-omni, an open-weight decision model that accepts audio, video, image, and text input in a single API call. Clef-flash's price falls from $0.09 to $0.038 per M input tokens, while its hosted context window drops from 64k to 24k. Cloudflare also reports median latency reductions of 1.7 to 2.0 times for the Clef model on Workers AI.

  9. Perplexity DevelopersOfficialAI score34

    Perplexity releases cookbook for a browser agent using the Decisions API

    AIPerplexity Developers says its new cookbook builds a browser agent that sends a screenshot and questions to pplx-decider-v1.1-27b through the Decisions API, which accepts text and image inputs. The developer's code converts the returned probabilities into clicks, scrolls, and stops.

  10. Tencent · new models on Hugging FaceOfficialAI score41

    Tencent Releases Youtu-Parsing-Omni, a 5B Omni-Modal Document and Media Parsing Model

    AITencent has open-sourced Youtu-Parsing-Omni, a 5B-parameter omni-modal model that outputs a single structured JSON covering layout, text, tables, formulas, ASR, OCR, and video segments. It scores 96.96 Overall on OmniDocBench, the highest among the compared models, and ships with weights on Hugging Face, a vLLM plugin, and inference examples.

  11. IThome · AINewsAI score55

    Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark

    AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.

  12. PandailyNewsAI score45

    Doubao Work Adds Infinite Creation Canvas, Seedream 5.0 Flash and Doubao 2.1 Lite

    AIByteDance's Doubao Work has added an infinite creation canvas that places source materials, design options and finished output on one page, wired to the new Seedream 5.0 Flash image model. The update also adds Doubao 2.1 Lite, a lighter model aimed at everyday office tasks such as documents, spreadsheets and slide decks, with faster responses and lower credit consumption. The announcement included no benchmark results for either model.

  13. ModelScopeOfficialAI score63

    Google releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device search

    AIGoogle released EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0 for private, on-device search and retrieval. It maps text, code, images, video, and audio into one shared space and reports a 9.92-point gain over EmbeddingGemma 1 on MTEB Code. The post lists about 191MB active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro.

    Video from @ModelScope2022's post

Oct 8

Oct 8Thu
  1. Higgsfield AI 🧩OfficialAI score36

    Higgsfield Katana adds community presets for Claude video editing

    AIHiggsfield has released community presets for Higgsfield Katana, its AI video editing tool available inside Claude. Users can pick a preset for motion graphics, 3D animations, product launches, fashion, car, travel, or aura-farming edits, then add their own characters, products, or clothes to recreate it in Claude. More presets are coming soon.

    Video from @higgsfield's post
  2. QbitAINewsAI score62

    Google launches Gemini agent for office work, able to call Claude models

    AIGoogle Cloud introduced the Gemini agent, a general office agent that can search, write emails, build slides, analyze data, run code, and coordinate sub-agents. It can take on an enterprise identity with email, calendar, and account, and it selects underlying models automatically, including Anthropic's Claude. The article presents this alongside OpenAI's Dots and Meta's Muse as competing office and personal agents.

  3. LeiphoneNewsAI score58

    Alibaba's Qwen Roadmap Targets 5T to 10T Parameters Amid Self-Improving Model Work

    AIAt the Apsara Conference, Alibaba's Qwen team outlined a roadmap of Qwen4 followed by Qwen4.5 and Qwen5, aiming for 5T to 10T parameters. The article notes that Qwen3.8 reached 2.4T parameters and that Qwen3.8-Flash activates 6B parameters per inference while cutting training cost to one-ninth. It also describes Qwen3.8-Max running model-driven experiments in chip design and inference optimization, and multimodal updates including a video model slated for November.

  4. TechCrunch · AINewsAI score67

    ChatGPT adds Intelligent UI with interactive visuals alongside GPT-6

    AIOpenAI is rolling out Intelligent UI in ChatGPT, which integrates interactive visuals such as tappable buttons, charts, editable graphs, and custom calculators into conversations. The feature launches with GPT-6 for Pro, Plus, Business, and Enterprise users, arrives Thursday for free and Go tiers, and lets users reduce how many visuals appear.

  5. The Verge · AINewsAI score75

    ChatGPT's Intelligent UI adds interactive visuals and tools to answers

    AIOpenAI is rolling out Intelligent UI in ChatGPT, letting answers include diagrams, charts, forms, tappable buttons, and generated tools. The feature runs on GPT-6, which is rolling out to Plus, Pro, Business, and Enterprise users now and to Go and free tiers on Thursday, with Sol and Luna models assigned by tier.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  6. Xiaomi MiMoOfficialAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    AIXiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    Why it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  7. Sundar PichaiXAI score65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

    Video from @sundarpichai's post
  8. NVIDIA BlogOfficialAI score49

    How Developers Use Frontier AI Agents to Build Omniverse Simulations

    AIDevelopers are pairing frontier AI models with NVIDIA Omniverse libraries to turn simulation ideas into working applications, from humanoid warehouse simulators to autonomous-driving test environments. In the examples, developers direct AI agents through natural-language instructions and review results, while Omniverse provides GPU-accelerated physics, rendering and sensor simulation. One experiment reported a simulated Unitree G1 humanoid clearing a hurdle in 64 of 100 trials.

  9. Databricks BlogOfficialAI score29

    Biomedical Imaging's Real Bottleneck Is Data Access, Not AI Models

    AIHospitals, academic centers, medtech firms, and pharma companies all face the same obstacle: imaging data is locked in clinical systems and hard to share. The EXAM study across 20 institutions showed federated learning, which shares model weights rather than patient data, improved AUC by 16% on average. Collaboration remains difficult due to scanner and protocol heterogeneity, privacy governance, and the lack of a common data substrate.

  10. Higgsfield AI 🧩OfficialAI score34

    Higgsfield launches Katana AI video editing tool inside Claude

    AIHiggsfield introduced Katana, its most powerful AI video editing tool, powered by Claude Motion and available now inside Claude via Higgsfield MCP. Users can upload a reference to create editable motion graphics, product launch videos, or aura-farming edits.

    Video from @higgsfield's post
  11. RunwayOfficialAI score36

    Runway promotes Claude Motion for animating charts and explainers

    AIRunway's post promotes using Claude Motion to animate charts, create customer walkthroughs, and make short explainers. The animated work can then be brought into Runway to generate videos and images. Claude Motion is described as being in beta, per the quoted Claude post.

    Video from @runwayml's post
  12. The DecoderNewsAI score65

    Anthropic launches Claude Dashboards and Motion features in beta

    AIAnthropic launched two beta features for Claude: Dashboards, which turns connected data sources like BigQuery, Databricks, Snowflake, or Salesforce into auto-updating live dashboards from text prompts, and Motion, which creates animated explainer videos from text, diagrams, and images. Dashboards is available to paid users and Motion to Team and Enterprise plans, while Docs, Slides, and Design leave beta and work across all plans, including free accounts.

  13. Google ResearchOfficialAI score22

    Google Research livestreams EmbeddingGemma 2 demo at COLM 2026 today

    AIGoogle Research is hosting a live demonstration of EmbeddingGemma 2 at its COLM booth #107 today at 12:00pm. The open multimodal model unifies text, image, audio, and video representations, with Sahil Dua available to connect with attendees.

    Video from @GoogleResearch's post
  14. OdysseyOfficialAI score34

    Odyssey-3 world knowledge can be applied to physical AI systems

    AIOdyssey says its Odyssey-3 model's learned world knowledge can be adapted by physical AI developers to control robots, power humanoids, drive cars, and fly drones. The post describes this as a capability for autonomous machines generally, without providing benchmarks, specifications, or availability details.

    Video from @odysseyml's post
  15. OdysseyOfficialAI score38

    Odyssey-3 is a foundation world model for physical AI and agents

    AIOdyssey announced Odyssey-3, a foundation world model it says enables applications in physical AI, human experiences, and training intelligences. The company highlights agents learning from experience inside Odyssey-3 while working toward objectives.

    Video from @odysseyml's post
  16. Google GemmaOfficialAI score27

    EmbeddingGemma 2 developer guide released by Google

    AIGoogle Gemma has published a developer guide for EmbeddingGemma 2, with code snippets to help developers start searching beyond text. The post directs readers to the full guide on the Google Developers Blog.

  17. Google GemmaOfficialAI score44

    Google publishes a developer guide for EmbeddingGemma 2 multimodal embeddings

    AIGoogle Gemma announces a developer guide showing how to embed text, code, images, video, audio, and interleaved inputs with EmbeddingGemma 2 using the sentence-transformers library. The guide outlines a four-step workflow: loading the model, embedding text and code with task prompts, embedding multimodal inputs, and optionally truncating dimensions with Matryoshka.

    Image from @googlegemma's post
  18. Nous ResearchOfficialAI score31

    Hermes Agent runs on ASUS ProArt RTX Spark PCs with local models

    AINous Research says Hermes Agent is now available on the new ASUS ProArt RTX Spark PCs, with MuseTree and ComfyUI integrations and local model support on up to 128 GB of unified memory. The post presents it as a creative stack that runs entirely on the user's own machine.