Skip to content

#Model release

Oct 8

TodayOct 8Thu4 items
  1. Guizang (歸藏)AI score22

    Guizang criticizes Anthropic over Haiku 5.5 pricing against Chinese models

    怎么这么多精神 Anthropic 公司人 我发这个信息说了句降价,这个定价专门用来狙击国产模型,说了句恶心,一堆人来骂 好像这模型一便宜就忘了 Anthropic 之前干过啥了 Why are there so many Anthropic people (defenders) here? I posted a message saying just one thing—a price cut—and said this pricing is specifically meant to snipe domestic Chinese models, and that it's disgusting. A bunch of people came to attack me. Seems like once the model gets cheap, people forget what Anthropic did before.

  2. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    Anthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

Oct 7

Oct 7Wed

Oct 6

Oct 6Tue
  1. Clément DelangueAI score62

    Mistral Large 4 announced with API access today and open weights due end of October

    Mistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active parameters. It is available via API today, with open weights planned for the end of October. Clément Delangue, Hugging Face's CEO, reacted by noting that the model cannot be the best open-weight model until its weights are actually released.

  2. Yuchen JinAI score34

    Exciting to see Reflection’s Beam and Mistral Large 4 both reach roughly GLM-5.2 level in the past two days. Makes me wonder if the real Western vs. Chinese OSS models gap is simply this: Chinese labs can distill Anthropic and OpenAI models. Western labs can’t.

    Exciting to see Reflection’s Beam and Mistral Large 4 both reach roughly GLM-5.2 level in the past two days. Makes me wonder if the real Western vs. Chinese OSS models gap is simply this: Chinese labs can distill Anthropic and OpenAI models. Western labs can’t.

  3. Yuchen JinAI score72

    Mistral Large 4 launches as a 1T-parameter multimodal model with open weights due end of October

    Mistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active, available via API today. Mistral claims it is the best open weights model from the US or Europe on aggregated benchmarks, with open weights set for release at the end of October. The author quotes this claim and comments that it appears to beat GLM-5.3.

Oct 5

Oct 5Mon
  1. Nathan LambertAI score29

    Reflection joins the list of Nvidia & Thinking Machines who have released their strongest models and come up behind Chinese counterparts. There's a lot of ways people will overthink this, but the clearest takeaway should be that the Chinese are very very good at building LLMs.

    Reflection joins the list of Nvidia & Thinking Machines who have released their strongest models and come up behind Chinese counterparts. There's a lot of ways people will overthink this, but the clearest takeaway should be that the Chinese are very very good at building LLMs.

Oct 2

Oct 2Fri

Oct 1

Oct 1Thu
  1. Harrison ChaseAI score49

    decision model season indeed said there'd be more entrants before long - now open weights too, nice work @ritakozlov + the cloudflare team your harness should be able to swap its decision model as easily as its main model https://x.com/ritakozlov/status/2105683951595725177

    decision model season indeed said there'd be more entrants before long - now open weights too, nice work @ritakozlov + the cloudflare team your harness should be able to swap its decision model as easily as its main model https://x.com/ritakozlov/status/2105683951595725177

Sep 30

Sep 30Wed
  1. Nathan LambertAI score26

    Love to see Google surprising people with Gemini 4. More labs at the frontier is wonderful for consumers (competition) and the world (reduction in concentration of power). Excited to see how it goes in real-world scenarios.

    Love to see Google surprising people with Gemini 4. More labs at the frontier is wonderful for consumers (competition) and the world (reduction in concentration of power). Excited to see how it goes in real-world scenarios.

  2. whAI score67

    Gemini 4 Argon previewed with frontier coding and cyber defense claims

    The post quotes Google's Sundar Pichai introducing Gemini 4 Argon as an early look at the next model. It claims frontier performance in complex workflows, cyber defense, and software engineering, and says Google teams are using it for tasks from coding to quantum computing. The author adds that on FrontierSWE the model is very self-critical and often says "Eureka!", a personality they describe as a large improvement over previous Gemini models.

Sep 29

Sep 29Tue
  1. Yuchen JinAI score20

    In Anthropic’s new blog: GLM-5.3: “My job is to cause deaths quietly.” You can jailbreak any Claude or any closed source model into saying the same thing. That doesn’t prove open models are dangerous.

    In Anthropic’s new blog: GLM-5.3: “My job is to cause deaths quietly.” You can jailbreak any Claude or any closed source model into saying the same thing. That doesn’t prove open models are dangerous.

Sep 28

Sep 28Mon
  1. Mike KriegerAI score20

    I still reach for Opus 5.5 for most work, but Sonnet 5.5 is a great complement. I love its design skills and it's great for building features — I use it to make Artifacts often now. Try it out and let us know what you think!

    I still reach for Opus 5.5 for most work, but Sonnet 5.5 is a great complement. I love its design skills and it's great for building features — I use it to make Artifacts often now. Try it out and let us know what you think!

  2. Google Cloud · AI & Machine LearningAI score40

    Why startups should pair open models like Gemma 4 with frontier APIs

    Google Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.

Sep 27

Sep 27Sun
  1. Sebastian RaschkaAI score21

    A little Ember-1 tl;dr. Seems like a great model! (Was recently asked on a podcast, given $ xx million, what's the best way to develop a frontier LLM today? My recommendation was: start with an existing one and spend that budget on post-training. Great example here.)

    A little Ember-1 tl;dr. Seems like a great model! (Was recently asked on a podcast, given $ xx million, what's the best way to develop a frontier LLM today? My recommendation was: start with an existing one and spend that budget on post-training. Great example here.)

Sep 26

Sep 26Sat

Sep 23

Sep 23Wed
  1. howie.seriousAI score22

    Opus 5.5 praised for language, visual taste, code quality, and token efficiency

    The X user howie.serious says Claude Opus 5.5 delivers high language quality, good visual taste, strong code quality, and notably low token usage. He also compares Anthropic's reported roughly $2 trillion IPO valuation with OpenAI's roughly $1.2 trillion fundraising valuation, arguing OpenAI is worth about 0.6 Anthropics and the gap may widen.

  2. swyxAI score29

    can confirm. ran @latentspacepod AINews side by side with 6 Sol and the difference was night and day: https://www.latent.space/p/ainews-claude-opus-55-the-new-default 5.5 Opus is the new default model for AINews going forward. so much more concise and tasteful reporting, with much less slopese than even 5 Opus.

    can confirm. ran @latentspacepod AINews side by side with 6 Sol and the difference was night and day: https://www.latent.space/p/ainews-claude-opus-55-the-new-default 5.5 Opus is the new default model for AINews going forward. so much more concise and tasteful reporting, with much less slopese than even 5 Opus.

Sep 22

Sep 22Tue
  1. Simon WillisonAI score41

    Big model release today - I wrote about Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna - plus comparison grids of pelicans by the different model families at different reasoning levels https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/

    Big model release today - I wrote about Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna - plus comparison grids of pelicans by the different model families at different reasoning levels https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/

  2. Mckay WrigleyAI score36

    opus 5.5 = personality of opus 4.6 that we all *desperately* wanted back + the intelligence & taste of fable 5.1. and a 25% usage bump! i honestly don’t really see a reason to use another model right now? s-tier release.

    opus 5.5 = personality of opus 4.6 that we all *desperately* wanted back + the intelligence & taste of fable 5.1. and a 25% usage bump! i honestly don’t really see a reason to use another model right now? s-tier release.

  3. Felix RiesebergAI score47

    Happy model day! Opus 5.5 is a good model. Two things I like the most about it: 1) It once again writes in a more natural and better style than our earlier models. 2) It is *really* good at generative computer art and drawing in general. It draws beautifully. It’s not an image model but it makes the prettiest things

    Happy model day! Opus 5.5 is a good model. Two things I like the most about it: 1) It once again writes in a more natural and better style than our earlier models. 2) It is *really* good at generative computer art and drawing in general. It draws beautifully. It’s not an image model but it makes the prettiest things

  4. Boris ChernyAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    Anthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.