Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Sep 22

Sep 22Tue
  1. Black Forest Labs · new models on Hugging FaceOfficialAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  2. AI SupremacyBlogAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

  3. Tencent HyOfficialAI score8

    Tencent Hunyuan confirms its model can edit images

    AITencent Hunyuan replied that its model can edit images, not just generate them, answering a user's question. The post names no specific model, version, or benchmark.

  4. X.PINXAI score46

    Moonshot's Kimi K3 now available on Amazon Bedrock

    AIMoonshot's Kimi K3 is now available on Amazon Bedrock, with its license requiring a paid agreement for model-hosting businesses and affiliates above $20M in annual revenue. AWS says customer data stays within its cloud, is not shared with Moonshot or used for training, and inference requests have zero data retention. Neither company disclosed financial terms.

    Image from @thexpin's post
  5. METR BlogOfficialAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

  6. Gemini API ChangelogOfficialAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. Kimi.aiOfficialAI score34

    Kimi K3 now available on Amazon Bedrock

    AIMoonshot AI's Kimi K3 is now available on Amazon Bedrock for coding, document analysis, and extended agent workflows. Bedrock provides access, encryption, and auditing controls, and explicit prompt caching is supported.

    Image from @Kimi_Moonshot's post
  2. StepFunOfficialAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  3. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  4. Claude Apps Release NotesOfficialAI score62

    Anthropic launches Claude Opus 5.5, first model in its 5.5 family

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    Why it matters: The release note gives a direct comparison to Claude Fable 5.1 and a 40% running cost reduction versus Opus 5, which helps readers weigh the tradeoff.

  5. Together AI BlogOfficialAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.

  6. VercelOfficialAI score26

    Vercel offers 40% off AI Gateway for Grok 4.7 through September 27

    AIVercel is offering 40% off AI Gateway usage for Grok 4.7 until September 27, with the discount applying to projects built with v0, Eve, and fx. The post promotes building with Grok 4.7 on Vercel at reduced cost and links to the changelog for details.

  7. Xiaomi MiMoOfficialAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  8. Xiaomi MiMoOfficialAI score42

    Xiaomi launches MiMo-V2.6 Pro UltraSpeed with up to 20× faster generation

    AIXiaomi's MiMo-V2.6-Pro UltraSpeed is now available in MiMo Desktop and through the API, delivering up to 20× faster generation. MiMo Desktop and membership plans launch alongside Pro and Flash, and users can run both models via subscription or their own API key. API pricing remains unchanged from V2.5.

  9. Xiaomi MiMoOfficialAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  10. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

    Video from @XiaomiMiMo's post
  11. Xiaomi MiMoOfficialAI score38

    Xiaomi MiMo-V2.6 raises intelligence at unchanged API prices

    AIXiaomi says MiMo-V2.6 keeps API pricing unchanged from V2.5 for both the Pro and Flash models while adding more intelligence. It claims MiMo-V2.6-Pro sets a new price-performance record among Chinese models, with comparable intelligence costing about 1/20 to 1/60 of leading international models.

    Image from @XiaomiMiMo's post
  12. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  13. Google AI DevelopersOfficialAI score8

    Google announces Gemini 3.8 Live with extended thinking

    AIGoogle AI Developers pointed readers to a Google blog post about Gemini 3.8 Live and its extended thinking capability. The post itself contains no further details, so specifics such as benchmarks, pricing, or availability cannot be confirmed from this source.

  14. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  15. Apple · new models on Hugging FaceOfficialAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.

  16. Xiaomi MiMoOfficialAI score34

    Xiaomi's MiMo Gallery showcases outputs generated entirely by MiMo-V2.6

    AIXiaomi's MiMo team released MiMo Gallery, a showcase where every 3D model, game, slide deck, video, music piece, and image was generated by MiMo-V2.6. The post previews MiMo-V2.6 as an upcoming model and points readers to the gallery for a first look.

    Video from @XiaomiMiMo's post
  17. NVIDIAOfficialAI score38

    Grok 4.7 launches as xAI's most capable model for coding

    AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.

  18. Xiaomi MiMo · new models on Hugging FaceOfficialAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  19. Xiaomi MiMo · new models on Hugging FaceOfficialAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  20. WorkBuddyOfficialAI score26

    WorkBuddy adds GLM-5.3-Flash to its model lineup

    AIWorkBuddy has added GLM-5.3-Flash to its model lineup and made it available now on the platform. The post invites users to try the model on their next task, but gives no further details on capabilities, speed, pricing, or benchmarks.

    Image from @WorkBuddy_AI's post

Sep 20

Sep 20Sun
  1. swyxXAI score22

    Jev Podcast Episode Announced by Latent Space Host swyx

    AIswyx announced a Latent Space podcast episode featuring Jev, subscribable on Apple and YouTube, and thanked guests Allen Park and Ke. A quoted post from @CompleteSkeptic claims Jev is a frontier model with 20-200x faster speed and 40-400x lower cost, but this post itself adds no verified details.

    Image from @swyx's post
  2. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

  3. QwenOfficialAI score56

    Qwen-Image-2.1 releases open weights for image generation and editing

    AIAlibaba's Qwen team released Qwen-Image-2.1 as an open-weights image model for both generation and editing, with a lightweight 7B architecture. The model natively generates and edits RGBA layers, supports up to 10 reference images for editing, and is available on GitHub, ModelScope, and Hugging Face.

    Image from @Alibaba_Qwen's post
  4. ModelScopeOfficialAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

    Why it matters: The post names concrete capabilities and a 7B size, letting readers compare it against the larger image models in the accompanying chart.

    Image from @ModelScope2022's post
  5. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    Why it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.