Skip to content

#Model release

Oct 8

TodayOct 8Thu70 items
  1. MarkTechPost58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  2. Leandro von Werra70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    Why it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  3. Thomas Wolf62

    Carbon-A open model finds 566 million candidate genes across 22,617 species

    The team released Carbon-A, an open model that finds genes directly in DNA, along with a database of 566.34 million candidate genes across 22,617 species. The model reads genomes without needing a close relative, and wet-lab validation in cats, chickens, and arabidopsis is cited, with 239 genes found missing from reference annotations of common species. The authors say the model marks gene locations but does not design DNA or predict gene function.

  4. Elvis Saravia32

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

  5. 雷峰网 Leiphone62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  6. Generalist14

    Other tasks our models can do: https://www.youtube.com/playlist?list=PLIv5TE4TO-ZgiMiHpD5ZSiJVLa9I6557R Read more about GEN-1.5, our latest foundation model for the physical world: https://generalistai.com/blog/gen-1.5

    Other tasks our models can do: https://www.youtube.com/playlist?list=PLIv5TE4TO-ZgiMiHpD5ZSiJVLa9I6557R Read more about GEN-1.5, our latest foundation model for the physical world: https://generalistai.com/blog/gen-1.5

  7. OpenRouter62

    StepFun's Step 5 Preview model is now available on OpenRouter at $1.00 per million input tokens

    OpenRouter announces that StepFun's Step 5 Preview is live on its platform, priced at $1.00 per million input tokens and $2.70 per million output tokens. Cache hits are 50% off at launch, bringing them to $0.05 per million. A week of free access is rolling out across partners including opencode, Cline, Nous Research, and Kilo Code.

  8. OpenRouter44

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

  9. Understanding AI (Timothy B. Lee)67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    TypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  10. StepFun18

    Want to learn more about Step 5 Preview? Model overview, benchmarks & demos ↓ https://www.stepfun.com/step-5-preview Technical specs & API docs ↓ https://platform.stepfun.ai/docs/en/guides/models/step-5-preview

    Want to learn more about Step 5 Preview? Model overview, benchmarks & demos ↓ https://www.stepfun.com/step-5-preview Technical specs & API docs ↓ https://platform.stepfun.ai/docs/en/guides/models/step-5-preview

  11. JetBrains AI Blog62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  12. The Decoder72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    Anthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

  13. 歸藏(guizang.ai)22

    Guizang criticizes Anthropic over Haiku 5.5 pricing against Chinese models

    怎么这么多精神 Anthropic 公司人 我发这个信息说了句降价,这个定价专门用来狙击国产模型,说了句恶心,一堆人来骂 好像这模型一便宜就忘了 Anthropic 之前干过啥了 Why are there so many Anthropic people (defenders) here? I posted a message saying just one thing—a price cut—and said this pricing is specifically meant to snipe domestic Chinese models, and that it's disgusting. A bunch of people came to attack me. Seems like once the model gets cheap, people forget what Anthropic did before.

  14. Testing Catalog58

    Daily AI brief covers Mistral, Google, OpenAI, Anthropic, Microsoft and other vendor news for October 8

    This daily brief from TestingCatalog collects recent AI announcements from several vendors, including Mistral Large 4, Claude Haiku 5.5, and GitHub stacked pull requests becoming generally available. The author states the brief was composed with Grok and cherry-picked news with post-editing. Items are mostly short product and pricing notices rather than detailed reporting, and several claims are unverified announcements.

  15. meng shao39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    Anthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

Oct 7

Oct 7Wed
  1. 极客公园 GeekPark46

    ChatGPT Adds Intelligent UI That Generates Interactive Tools, Google Launches Playground

    OpenAI said on October 7 that ChatGPT's new Intelligent UI will automatically combine text, charts, buttons and forms into interactive interfaces such as calculators and mini-games, rolling out to Plus, Pro, Business and Enterprise users from October 7 and to Free and Go users from October 8. Google also launched Playground, an experimental platform where users create, modify and play browser games from natural-language descriptions, initially for U.S. users aged 18 and older.

  2. IT之家 · 人工智能72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    Anthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  3. Databricks36

    Claude Haiku 5.5 launches on Databricks as a Day 0 release

    Anthropic's Claude Haiku 5.5 is available on Databricks from day zero, which Databricks calls its cheapest, fastest, and most capable small model. On Databricks' OfficeQA Pro V1 benchmark, it delivers about 15% higher quality than Haiku 4.5 at a fraction of the cost. Users can run it alongside 60+ other models on data already in Databricks, with Unity Gateway handling governance, monitoring, and security.