Skip to contentSkip to stories
Updated

#Multimodal

Oct 9

TodayOct 9Fri
  1. elvisXAI score40

    Syren Video learns your style to build AI videos from prompts

    AISyren Video, a new agentic video tool, learns preferred graphics, motion, and editing rhythm from a user's library and generates new videos from a prompt. Users refine the results through chat, and the tool is free to try in a browser or through Claude MCP, per the company's announcement. The post's author says the education sector is exploring it.

Oct 8

Oct 8Thu
  1. Alexander DoriaXAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post

Oct 3

Oct 3Sat

Oct 1

Oct 1Thu
  1. ReplicateOfficialAI score46

    Replicate powers Tavus's Griffin, a video Turing test-passing model

    AIReplicate says it is powering Griffin from Tavus on its platform. Tavus describes Griffin as the first model to pass the video Turing test, with 48% of live conversation participants believing it was a real human. The model ranks first on NVIDIA's full-duplex AI video benchmark.

Sep 30

Sep 30Wed
  1. ModelScopeOfficialAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Why it matters: The post gives specific benchmark, speed, and input details, so readers can judge how the model handles real-time scene exploration from images or video.

    Video from @ModelScope2022's post

Sep 28

Sep 28Mon
  1. ModelScopeOfficialAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post

Sep 17

Sep 17Thu
That’s everything