Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Sep 28

Sep 28Mon
  1. SpaceXAIOfficialAI score46

    Grok 4.7 is now available on Amazon Bedrock

    AIaccording to the SpaceXAI account, which is operated by xAI and Grok. The post gives no further details on pricing, context length, or benchmarks.

    Video from @SpaceXAI's post
  2. Simon WillisonXAI score55

    Sonnet 5.5 becomes the free-tier model on claude.ai

    AISimon Willison says Claude Sonnet 5.5 now powers the free tier on claude.ai, so free users can run the kinds of experiments he describes. He contrasts this with ChatGPT's free tier, which he says still runs the less capable GPT-5.6 Luna.

  3. Ali GhodsiXAI score62

    Databricks finds Opus 5.5 cheaper and better, GPT-6 Luna 20x cheaper per task

    AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.

  4. Mike KriegerXAI score67

    Anthropic releases Claude Sonnet 5.5, 30% faster and up to 30% cheaper than Sonnet 5

    AIAnthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family. The company says it is more than 30% faster than Sonnet 5 and costs up to 30% less for most work.

    Why it matters: The post gives concrete speed and price changes against Sonnet 5, which helps readers judge whether the upgrade fits their workloads and budgets.

  5. catXAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

    Video from @_catwu's post
  6. Boris ChernyXAI score38

    Claude Sonnet 5.5 runs 30% faster with 30% less usage

    AIAnthropic's Claude Sonnet 5.5, the second model in the Claude 5.5 family, is shown fixing a bug in Claude Code. Boris Cherny says it runs 30% faster and uses 30% less usage, and Anthropic's announcement says it runs over 30% faster and costs up to 30% less for most work.

    Video from @bcherny's post
  7. Alex AlbertXAI score62

    Claude Sonnet 5.5 Is Faster and Cheaper Than Sonnet 5, Per Anthropic

    AIAnthropic has introduced Claude Sonnet 5.5, the second model in the Claude 5.5 family, as a clear upgrade over Sonnet 5. The announcement says it runs more than 30% faster and costs up to 30% less for most work. Alex Albert, quoting the announcement, says the model writes clearly, is very fast, and makes a major capabilities jump over Sonnet 5.

  8. v0OfficialAI score62

    Claude Sonnet 5.5 is now available in v0

    AIwhich links to a page for trying the model. The quoted announcement describes it as the second model in the Claude 5.5 family, a clear upgrade over Sonnet 5 that runs more than 30% faster and costs up to 30% less for most work.

  9. AnthropicOfficialAI score61

    Claude Sonnet 5.5 is now available from Anthropic

    AIAnthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family. The quoted announcement says it is a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  10. ReplicateOfficialAI score28

    Pruna's P-Video-2-Pro video model now runs on Replicate

    AIReplicate has added P-Video-2-Pro, the latest video model from Pruna AI, which sits on the edge of the preference-speed and preference-price Pareto frontiers. Design Arena ranks its Quality and Speed variants tied for #2 on the Image to Video leaderboard with an Elo of 1325, with the Quality version generating in 8.0 seconds and the Speed version in 4.5 seconds.

  11. Google Cloud · AI & Machine LearningOfficialAI score40

    Why startups should pair open models like Gemma 4 with frontier APIs

    AIGoogle Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.

  12. AI at MetaOfficialAI score16

    Meta lists ten Muse models and products released over six months

    AIMeta's AI account says it has released a series of Muse-branded models and products in the past six months, listing Muse Spark, Muse Image, Muse Video, Muse Spark 1.1, Meta Model API, Muse Glimmer, Muse Spark 1.2, Muse Code, Muse Spark 1.3, and Muse. The post says the company is "just getting started" without giving specifications, benchmarks, or pricing.

    Video from @AIatMeta's post
  13. ModelScopeOfficialAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post

Sep 27

Sep 27Sun
  1. KhazixXAI score10

    Khazix says Claude Opus 5.5 is exceptionally strong

    AIKhazix (@Khazix0918) says Claude Opus 5.5 is exceptionally strong, to the point of providing a full day's worth of laughs. The post gives no benchmark scores, pricing, or further technical details.

    Image from @Khazix0918's post
  2. Claude Apps Release NotesOfficialAI score65

    Anthropic launches Claude Sonnet 5.5 as second Claude 5.5 model

    AIAnthropic has launched Claude Sonnet 5.5, the second model in its Claude 5.5 family. The company describes it as a faster, lower-cost complement to Claude Opus 5.5, and points readers to a blog post for more information.

    Why it matters: The release note places Sonnet 5.5 beside Opus 5.5 in the Claude 5.5 family, clarifying which model suits speed and cost needs.

  3. Amp NewsOfficialAI score67

    Amp switches its default medium mode to Claude Opus 5.5

    AIAmp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.

    Why it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.

  4. Sebastian RaschkaXAI score21

    Ember-1 is Kimi K3 post-trained for 40% more concise reasoning

    AIEmber-1, a model built on Kimi K3 through post-training, reportedly reasons 40% more concisely while keeping the same quality and running 40% faster and cheaper. Sebastian Raschka cites it as an example that starting from an existing frontier LLM and investing the budget in post-training is an effective development path.

    Image from @rasbt's post
  5. Felix RiesebergXAI score13

    Felix Rieseberg Rebuilds His Homepage Using Opus 5.5

    AIFelix Rieseberg, an Anthropic employee, says he remade his homepage with Opus 5.5 and pushed it hard, using it to create music, movies, textures, and Blender models. He says he is very happy with the result and links to his site.

    Image from @felixrieseberg's post
  6. MiniMax (official)OfficialAI score34

    MiniMax-M3.1 Flash Preview launches on Token Plan for high-volume teams

    AIMiniMax has made MiniMax-M3.1 Flash Preview available on its Token Plan, targeting teams with high-volume, latency-sensitive workloads. The model is faster and lighter, and it can be used under an existing Token Plan subscription without extra setup. MiniMax also says the text model, M3.1-Flash-Preview, debuted on MiniMax Code for everyday development tasks.

  7. Tibor BlahoXAI score85

    OpenAI releases GPT-6 Sol and Luna as Anthropic launches Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.

    Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.

    Video from @btibor91's post
  8. Tibor BlahoXAI score71

    OpenAI and Anthropic ship GPT-6 Sol and Luna and Claude Opus 5.5 in the same week

    AIOpenAI released GPT-6 Sol and Luna at API prices 50% below GPT-5.6 promotional pricing, and Anthropic released Claude Opus 5.5 the same day at 40% less than Opus 5. The roundup also covers Claude Code cloud sessions reaching general availability, the Claude Marketplace launch, OpenAI's new misalignment disclosures after the Hugging Face incident, and DevDay on September 29. The post is a relayed weekly digest, and it includes the author's closing promotion for AIPRM, which is not part of the reported news.

  9. The SequenceBlogAI score55

    Opus 5.5 cuts costs while Meta and US–China talks widen AI's reach

    AIAnthropic released Claude Opus 5.5 at about 40% lower cost than Opus 5, priced at $4/$20 per MTok input/output. Meta said Muse is coming to its AI glasses in the coming months, while Washington and Beijing held their first AI dialogue and discussed an incident-notification channel. The newsletter argues that costs, interfaces, experiments, and diplomacy increasingly determine how much value AI creates.

  10. Xiaomi MiMo · new models on Hugging FaceOfficialAI score44

    Xiaomi releases MiMo-V2.6-Flash-MOPD, an upgraded MoE model with 1M context

    AIXiaomi has released MiMo-V2.6-Flash-MOPD on Hugging Face, an upgrade of the MiMo-V2.6-Flash-RL checkpoint that fuses several domain-specialized teachers into one model. The sparse MoE model has 309B total and 15B activated parameters, a 1M-token context length, and supports text, image, video, and audio inputs. The checkpoint targets tool-call repetition, a failure mode where the model repeatedly issues the same or similar tool calls without making progress.

  11. Xiaomi MiMoOfficialAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi releases MiMo-V2.6-Pro-MOPD, a 1.02T-parameter sparse MoE model

    AIXiaomi has released MiMo-V2.6-Pro-MOPD, an upgrade of the MiMo-V2.6-Pro-RL checkpoint that fuses several domain-specialized teachers into one model via MOPD2 and targets tool-call repetition. The sparse MoE model has 1.02T total and 42B activated parameters, a 1M-token context length, and accepts text, image, video, and audio inputs. Weights are available on Hugging Face and ModelScope, with deployment recipes for SGLang and vLLM.

  2. KhazixXAI score18

    Khazix says Claude Opus 5.5 excels at long-running agent tasks

    AIThe author ran a full-rewrite-scale task with Claude Opus 5.5 for over six hours, reading every thinking summary along the way. They call it stable and strong, rating it top tier across communication, comprehension, aesthetics, development, and long-horizon agent work, and wish OpenAI would catch up.

    Image from @Khazix0918's post
  3. OpenCodeOfficialAI score34

    LongCat-2.5-Preview Free on OpenCode for Two Weeks

    AILongCat-2.5-Preview is free on OpenCode for two weeks, offering a 1M context window, multimodal support, and zero data retention. The post does not provide further details on pricing terms or capabilities beyond these listed features.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  6. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. Google AIOfficialAI score57

    Google AI lists weekly releases including Gemini 3.8 TTS, Live Avatar, and Project Suncatcher

    AIGoogle AI's weekly roundup lists Gemini 3.8 Flash TTS and Flash-Lite TTS as expressive audio generation models. It also announces Gemini 3.8 Live with Live Avatar for near real-time visual conversation and a Live Chat voice feature on the Gemini Notebook mobile app across about 100 languages. Project Suncatcher will launch a prototype satellite to test Google TPUs in orbit and explore solar-powered AI compute in space.

  2. WorkBuddyOfficialAI score36

    Grok-4.7 now available on WorkBuddy for user tasks

    AIWorkBuddy has added Grok-4.7 to its model lineup, making it available for users to try on their next task. The post gives no benchmark, pricing, context length, or other technical details.

    Image from @WorkBuddy_AI's post
  3. NewcomerBlogAI score46

    Meta's Muse and Instinct lead a race to build consumer AI agents

    AIMeta has launched Muse, a personal AI assistant that can check out across shops on Shopify, with Shopify stock rising more than 10% on the news. Startup Instinct is reportedly seeking new funding at a $10 billion valuation, up from a round valuing it at $2.5 billion, while Amazon has said it will block Muse from shopping on its platform.