Skip to content

All AI news

Oct 8

TodayOct 8Thu55 items
  1. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    Shanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  2. QbitAI (量子位)AI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    Anthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  3. Xiaomi MiMoAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    Xiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    AIWhy it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  4. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  5. Artificial AnalysisAI score7

    Artificial Analysis publishes AA-Video-T2V v2.0 prompt for snowy cabin scene

    Artificial Analysis shares the second part of an AA-Video-T2V v2.0 prompt describing a four-shot documentary-style handheld video of a glass cabin in falling snow. The shots follow a caretaker sweeping snow from the deck, empty snow-covered windows, an empty interior, and the same caretaker stamping snow off his boots at the door, with hard cuts between shots.

  6. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

  7. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

  8. Artificial AnalysisAI score46

    Grok Imagine Video 1.5 Lite ranks ahead of Google's Veo 3.1 on AA-Video-T2V v2.0, at about a third of the price At 1080p with audio, Grok Imagine Video 1.5 Lite costs $0.14 per second, against $0.40 per second for Veo 3.1. It ranks #17 on AA-Video-T2V v2.0, two places above Veo 3.1. Against Grok Imagine Video 1.5, Lite costs 44% less at 1080p and ranks six places lower.

    Grok Imagine Video 1.5 Lite ranks ahead of Google's Veo 3.1 on AA-Video-T2V v2.0, at about a third of the price At 1080p with audio, Grok Imagine Video 1.5 Lite costs $0.14 per second, against $0.40 per second for Veo 3.1. It ranks #17 on AA-Video-T2V v2.0, two places above Veo 3.1. Against Grok Imagine Video 1.5, Lite costs 44% less at 1080p and ranks six places lower.

  9. Artificial AnalysisAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    SpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

  10. ClaudeDevsAI score42

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

  11. Artificial AnalysisAI score38

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

  12. Artificial AnalysisAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).

  13. OpenAI DevelopersAI score38

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

  14. Elvis SaraviaAI score42

    Huge release from @odysseyml. Odyssey-3 Pro sets a new top score on Physics-IQ Verified, a benchmark that asks models to continue videos of real physics experiments. The robotics results stood out to me. With tens of hours of demos, the robot arm recovered from a missed grasp, a behavior that never appeared in those demos.

    Huge release from @odysseyml. Odyssey-3 Pro sets a new top score on Physics-IQ Verified, a benchmark that asks models to continue videos of real physics experiments. The robotics results stood out to me. With tens of hours of demos, the robot arm recovered from a missed grasp, a behavior that never appeared in those demos.

  15. SantiagoAI score46

    The #1 video-to-video model in the Physics-IQ Verified benchmark is finally live! (Their research preview is) Odyssey 3 Pro is a world model, and nobody beats it for physical accuracy. You can use this model to control a robot, drive a car, play a video game, or pilot a drone. • It takes visual observations from the world • Uses these observations to learn how things work • Then maps that knowledge to the system's physical controls

    The #1 video-to-video model in the Physics-IQ Verified benchmark is finally live! (Their research preview is) Odyssey 3 Pro is a world model, and nobody beats it for physical accuracy. You can use this model to control a robot, drive a car, play a video game, or pilot a drone. • It takes visual observations from the world • Uses these observations to learn how things work • Then maps that knowledge to the system's physical controls

  16. OdysseyAI score31

    We believe world models will power increasingly capable physical AI, generate environments to train intelligences, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this. Experience Odyssey-3 today! https://odyssey.systems/meet-odyssey-3

    We believe world models will power increasingly capable physical AI, generate environments to train intelligences, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this. Experience Odyssey-3 today! https://odyssey.systems/meet-odyssey-3

  17. OdysseyAI score34

    Odyssey-3’s learned world knowledge can also be applied to physical systems, enabling physical AI developers to adapt Odyssey-3 to control robots, power humanoids, drive cars, fly drones, and other autonomous machines.

    Odyssey-3’s learned world knowledge can also be applied to physical systems, enabling physical AI developers to adapt Odyssey-3 to control robots, power humanoids, drive cars, fly drones, and other autonomous machines.

  18. OdysseyAI score38

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

  19. OdysseyAI score22

    Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s video-to-video benchmark, achieving 66.1, the highest reported score. Physics-IQ tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

    Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s video-to-video benchmark, achieving 66.1, the highest reported score. Physics-IQ tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

  20. OdysseyAI score40

    Today we're launching Odyssey-3, the most powerful foundation world model yet. It sets a new state of the art on Physics-IQ, and powers robots, trains AIs, and generates interactive experiences. It's really cool. Experience the model today, all for free!

    Today we're launching Odyssey-3, the most powerful foundation world model yet. It sets a new state of the art on Physics-IQ, and powers robots, trains AIs, and generates interactive experiences. It's really cool. Experience the model today, all for free!

  21. OdysseyAI score42

    Odyssey-3 can generate interactive environments from a prompt, all in real time. It's a new kind of world simulator, and learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations.

    Odyssey-3 can generate interactive environments from a prompt, all in real time. It's a new kind of world simulator, and learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations.

  22. MarkTechPostAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  23. Leandro von WerraAI score70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    AIWhy it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  24. Thomas WolfAI score62

    Carbon-A open model finds 566 million candidate genes across 22,617 species

    The team released Carbon-A, an open model that finds genes directly in DNA, along with a database of 566.34 million candidate genes across 22,617 species. The model reads genomes without needing a close relative, and wet-lab validation in cats, chickens, and arabidopsis is cited, with 239 genes found missing from reference annotations of common species. The authors say the model marks gene locations but does not design DNA or predict gene function.

  25. Elvis SaraviaAI score32

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

  26. Leiphone (雷峰网)AI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.