Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 22

Sep 22Tue
  1. Sherwin WuXAI score46

    GPT-6 Luna launches at $0.10 and $0.50 per million tokens

    AIOpenAI's GPT-6 Luna is priced at $0.10 per 1M input tokens and $0.50 per 1M output tokens, according to Sherwin Wu. Per OpenAI Devs, Luna and GPT-6 Sol launch today with API prices 50% lower than GPT-5.6. Wu jokes that per-billion-token pricing may soon be needed.

  2. ChatGPTOfficialAI score62

    OpenAI rolls out GPT-6 Sol and GPT-6 Luna in ChatGPT Work and Codex

    AIOpenAI announced GPT-6 Sol and GPT-6 Luna, rolling out today in ChatGPT Work and Codex. The rollout covers Plus, Pro, Business, Enterprise, and Edu users.

    Why it matters: The post names two new GPT-6 variants and their rollout to specific ChatGPT and Codex plan tiers, which shows how access is being staged.

    Video from @ChatGPT's post
  3. ChatGPTOfficialAI score72

    GPT-6 Luna Rolls Out to Free and Go Users in the ChatGPT Desktop App

    AIOpenAI's official ChatGPT account says Free and Go users can try GPT-6 Luna in the desktop app, with rollout starting today. The post links to OpenAI's announcement introducing GPT-6 Sol and Luna.

    Why it matters: The post names the access tier and platform for the new model, which is the detail readers need to judge whether it applies to them.

  4. Felix RiesebergXAI score47

    Anthropic's Opus 5.5 praised for natural writing and computer art

    AIAnthropic's Felix Rieseberg says the new Opus 5.5 model writes more naturally than earlier models. He also highlights its strong generative computer art and drawing ability, noting it is not an image model yet produces attractive visuals.

  5. Mike KriegerXAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  6. Lydia Hallie ✨XAI score23

    Claude paid plans get a usage limit reset until Oct 22

    AIAnthropic's Pro, Max, and Team subscribers can claim a usage limit reset in Settings → Usage, available until Oct 22. Opus 5.5 is now the default for paid plans and is priced lower than Opus 5, so 5-hour and weekly limits go 25% further.

  7. v0OfficialAI score52

    Claude Opus 5.5 is now available in v0

    AIv0 announced that users can now use Claude Opus 5.5 in v0, with a link to try it. The quoted Anthropic post says Opus 5.5 is the first model in its new Claude 5.5 family, performs at the level of Claude Fable 5.1 on most tasks, and costs 40% less to run than Opus 5.

  8. Boris ChernyXAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

    Why it matters: The author reports a same-task comparison in which Opus 5.5 finished a HAProxy C-to-Rust port faster and cheaper than Fable 5.1, offering a concrete cost and time benchmark.

  9. Lydia Hallie ✨XAI score62

    Opus 5.5 released, claimed about 30% faster and 40% cheaper per task than Opus 5

    AIOpus 5.5 is now available, and the author says it is about 30% faster and about 40% cheaper per task than Opus 5. The author describes it as feeling like Fable and invites readers to try it.

    Why it matters: The post gives concrete speed and cost comparisons against Opus 5, the most useful part for judging whether the model fits a workload.

    Video from @lydiahallie's post
  10. Alex AlbertXAI score62

    Anthropic introduces Claude Opus 5.5 as first model in Claude 5.5 family

    AIAnthropic has introduced Claude Opus 5.5, the first model in its new Claude 5.5 family. The quoted announcement says it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. Alex Albert's post praises the model as smart, clear, fast, and cheaper, but offers no independent test results.

    Why it matters: The quoted announcement gives a concrete cost comparison against Opus 5, which helps readers weigh the model's value beyond the author's praise.

  11. catXAI score62

    Claude Opus 5.5 becomes the default model in Claude Code and Claude app

    AIClaude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. Anthropic is defaulting to effort medium across products, which it says is comparable to Fable 5.1 on intelligence but faster. Rate limits will go 25% further on Opus 5.5 compared to Opus 5.

    Why it matters: The post names concrete default changes across Claude Code and the Claude app, plus a specific effort setting and rate-limit difference, useful for judging day-to-day cost and speed.

  12. Sam BowmanXAI score75

    Anthropic's Sam Bowman says Claude Opus 5.5 is safer, reducing misalignment risk

    AISam Bowman says Claude Opus 5.5 is sufficiently safer than its predecessors that releasing it more likely than not reduces misalignment risks. The quoted @claudeai post introduces Claude Opus 5.5 as the first model in the Claude 5.5 family, performing at the level of Claude Fable 5.1 on most tasks at 40% lower run cost than Opus 5.

    Why it matters: The post links a safety judgment to a model release, which is useful for readers weighing how Anthropic frames release decisions against misalignment risk.

  13. AnthropicOfficialAI score71

    Anthropic releases Claude Opus 5.5, the first model in its Claude 5.5 family

    AIAnthropic has made Claude Opus 5.5 available today, introducing it as the first model in its new Claude 5.5 family. According to the quoted @claudeai post, it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

    Why it matters: The post gives a concrete cost comparison against Opus 5 and names the model family, helping readers gauge the trade-off between price and performance.

  14. Lovable BlogOfficialAI score38

    Lovable Adds Opus 5.5, Cutting Build Steps by a Third to Half at Same Quality

    AILovable now offers Opus 5.5, which it says matches Opus 5's results while finishing builds in a third to half fewer steps. Internal benchmarks showed Opus 5.5 scoring 4 to 6% ahead of Opus 5 on verification discipline, with step reductions of 26% to 57% and input token reductions of 21% to 59% across tasks.

  15. Black Forest Labs · new models on Hugging FaceOfficialAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  16. Black Forest Labs · new models on Hugging FaceOfficialAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

  17. Black Forest Labs · new models on Hugging FaceOfficialAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  18. AI SupremacyBlogAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

  19. X.PINXAI score46

    Moonshot's Kimi K3 now available on Amazon Bedrock

    AIMoonshot's Kimi K3 is now available on Amazon Bedrock, with its license requiring a paid agreement for model-hosting businesses and affiliates above $20M in annual revenue. AWS says customer data stays within its cloud, is not shared with Moonshot or used for training, and inference requests have zero data retention. Neither company disclosed financial terms.

    Image from @thexpin's post
  20. METR BlogOfficialAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

  21. Gemini API ChangelogOfficialAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. Kimi.aiOfficialAI score34

    Kimi K3 now available on Amazon Bedrock

    AIMoonshot AI's Kimi K3 is now available on Amazon Bedrock for coding, document analysis, and extended agent workflows. Bedrock provides access, encryption, and auditing controls, and explicit prompt caching is supported.

    Image from @Kimi_Moonshot's post
  2. StepFunOfficialAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  3. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  4. Claude Apps Release NotesOfficialAI score62

    Anthropic launches Claude Opus 5.5, first model in its 5.5 family

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    Why it matters: The release note gives a direct comparison to Claude Fable 5.1 and a 40% running cost reduction versus Opus 5, which helps readers weigh the tradeoff.

  5. Together AI BlogOfficialAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.

  6. VercelOfficialAI score26

    Vercel offers 40% off AI Gateway for Grok 4.7 through September 27

    AIVercel is offering 40% off AI Gateway usage for Grok 4.7 until September 27, with the discount applying to projects built with v0, Eve, and fx. The post promotes building with Grok 4.7 on Vercel at reduced cost and links to the changelog for details.

  7. Xiaomi MiMoOfficialAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  8. Xiaomi MiMoOfficialAI score42

    Xiaomi launches MiMo-V2.6 Pro UltraSpeed with up to 20× faster generation

    AIXiaomi's MiMo-V2.6-Pro UltraSpeed is now available in MiMo Desktop and through the API, delivering up to 20× faster generation. MiMo Desktop and membership plans launch alongside Pro and Flash, and users can run both models via subscription or their own API key. API pricing remains unchanged from V2.5.

  9. Xiaomi MiMoOfficialAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  10. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

    Video from @XiaomiMiMo's post
  11. Xiaomi MiMoOfficialAI score38

    Xiaomi MiMo-V2.6 raises intelligence at unchanged API prices

    AIXiaomi says MiMo-V2.6 keeps API pricing unchanged from V2.5 for both the Pro and Flash models while adding more intelligence. It claims MiMo-V2.6-Pro sets a new price-performance record among Chinese models, with comparable intelligence costing about 1/20 to 1/60 of leading international models.

    Image from @XiaomiMiMo's post
  12. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  13. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  14. Apple · new models on Hugging FaceOfficialAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.