Skip to content

#Model release

Oct 8

TodayOct 8Thu40 items
  1. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    Shanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  2. QbitAI (量子位)AI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    Anthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  3. Xiaomi MiMoAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    Xiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    AIWhy it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  4. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  5. OpenAI DevelopersAI score38

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

    API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts.

  6. Elvis SaraviaAI score42

    Huge release from @odysseyml. Odyssey-3 Pro sets a new top score on Physics-IQ Verified, a benchmark that asks models to continue videos of real physics experiments. The robotics results stood out to me. With tens of hours of demos, the robot arm recovered from a missed grasp, a behavior that never appeared in those demos.

    Huge release from @odysseyml. Odyssey-3 Pro sets a new top score on Physics-IQ Verified, a benchmark that asks models to continue videos of real physics experiments. The robotics results stood out to me. With tens of hours of demos, the robot arm recovered from a missed grasp, a behavior that never appeared in those demos.

  7. SantiagoAI score46

    The #1 video-to-video model in the Physics-IQ Verified benchmark is finally live! (Their research preview is) Odyssey 3 Pro is a world model, and nobody beats it for physical accuracy. You can use this model to control a robot, drive a car, play a video game, or pilot a drone. • It takes visual observations from the world • Uses these observations to learn how things work • Then maps that knowledge to the system's physical controls

    The #1 video-to-video model in the Physics-IQ Verified benchmark is finally live! (Their research preview is) Odyssey 3 Pro is a world model, and nobody beats it for physical accuracy. You can use this model to control a robot, drive a car, play a video game, or pilot a drone. • It takes visual observations from the world • Uses these observations to learn how things work • Then maps that knowledge to the system's physical controls

  8. OdysseyAI score31

    We believe world models will power increasingly capable physical AI, generate environments to train intelligences, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this. Experience Odyssey-3 today! https://odyssey.systems/meet-odyssey-3

    We believe world models will power increasingly capable physical AI, generate environments to train intelligences, and enable new kinds of human experiences, and we believe Odyssey-3 is a big leap towards this. Experience Odyssey-3 today! https://odyssey.systems/meet-odyssey-3

  9. OdysseyAI score34

    Odyssey-3’s learned world knowledge can also be applied to physical systems, enabling physical AI developers to adapt Odyssey-3 to control robots, power humanoids, drive cars, fly drones, and other autonomous machines.

    Odyssey-3’s learned world knowledge can also be applied to physical systems, enabling physical AI developers to adapt Odyssey-3 to control robots, power humanoids, drive cars, fly drones, and other autonomous machines.

  10. OdysseyAI score38

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

  11. OdysseyAI score40

    Today we're launching Odyssey-3, the most powerful foundation world model yet. It sets a new state of the art on Physics-IQ, and powers robots, trains AIs, and generates interactive experiences. It's really cool. Experience the model today, all for free!

    Today we're launching Odyssey-3, the most powerful foundation world model yet. It sets a new state of the art on Physics-IQ, and powers robots, trains AIs, and generates interactive experiences. It's really cool. Experience the model today, all for free!

  12. OdysseyAI score42

    Odyssey-3 can generate interactive environments from a prompt, all in real time. It's a new kind of world simulator, and learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations.

    Odyssey-3 can generate interactive environments from a prompt, all in real time. It's a new kind of world simulator, and learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations.

  13. MarkTechPostAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  14. Leandro von WerraAI score70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    AIWhy it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  15. Thomas WolfAI score62

    Carbon-A open model finds 566 million candidate genes across 22,617 species

    The team released Carbon-A, an open model that finds genes directly in DNA, along with a database of 566.34 million candidate genes across 22,617 species. The model reads genomes without needing a close relative, and wet-lab validation in cats, chickens, and arabidopsis is cited, with 239 genes found missing from reference annotations of common species. The authors say the model marks gene locations but does not design DNA or predict gene function.

  16. Elvis SaraviaAI score32

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

  17. Leiphone (雷峰网)AI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  18. GeneralistAI score14

    Other tasks our models can do: https://www.youtube.com/playlist?list=PLIv5TE4TO-ZgiMiHpD5ZSiJVLa9I6557R Read more about GEN-1.5, our latest foundation model for the physical world: https://generalistai.com/blog/gen-1.5

    Other tasks our models can do: https://www.youtube.com/playlist?list=PLIv5TE4TO-ZgiMiHpD5ZSiJVLa9I6557R Read more about GEN-1.5, our latest foundation model for the physical world: https://generalistai.com/blog/gen-1.5

  19. OpenRouterAI score44

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

  20. Understanding AI (Timothy B. Lee)AI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    TypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  21. StepFunAI score18

    Want to learn more about Step 5 Preview? Model overview, benchmarks & demos ↓ https://www.stepfun.com/step-5-preview Technical specs & API docs ↓ https://platform.stepfun.ai/docs/en/guides/models/step-5-preview

    Want to learn more about Step 5 Preview? Model overview, benchmarks & demos ↓ https://www.stepfun.com/step-5-preview Technical specs & API docs ↓ https://platform.stepfun.ai/docs/en/guides/models/step-5-preview

  22. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    AIWhy it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  23. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    Anthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.