Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 9

Jul 9Thu
  1. Meta AI BlogOfficialAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Aman SangerXAI score40

    Cursor and SpaceXAI release Grok 4.5, a model trained from scratch

    AICursor's Aman Sanger says the new model is an enormous improvement over Composer 2.5 and was trained entirely from scratch with the SpaceXAI team. The quoted Cursor post identifies it as Grok 4.5, Cursor's most powerful model yet and the first built for more than software engineering.

  2. Michael TruellXAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

Jul 7

Jul 7Tue
  1. Meta AI BlogOfficialAI score75

    Meta launches Muse Image, an agentic image model with search and code tools

    AIMeta Superintelligence Labs has released Muse Image, which can invoke search and coding tools and self-refine its generations before output. It is available today in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook coming soon. Meta also previewed Muse Video, which is coming soon to creators and Meta AI and is reported as ranking No. 3 on Arena for text-to-video at the time of writing.

    Why it matters: The source describes how search, code execution, and self-refinement change image generation, which matters to anyone comparing agentic media models with plain prompt-to-image systems.

Jul 3

Jul 3Fri
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score22

    XiaomiMiMo releases MiMo-V2.5-DFlash model weights on Hugging Face

    AIA model repository named XiaomiMiMo/MiMo-V2.5-DFlash is listed on Hugging Face with 311B parameters and tensor types F32, BF16, and F8_E4M3. The README is empty, and the page reports 434 downloads last month and no Inference Provider deployment.

Jul 1

Jul 1Wed
  1. Alex AlbertXAI score23

    Anthropic welcomes back Fable 5 model

    AIAnthropic's Alex Albert celebrated the return of Fable 5, with Claude's official account announcing its comeback. The post offers no details on capabilities, pricing, availability, or the reason for its return.

  2. Mistral AI · new models on Hugging FaceOfficialAI score54

    Mistral AI releases Leanstral 1.5, an open-source Lean 4 code agent model

    AIMistral AI released Leanstral 1.5 on Hugging Face as an open-source code agent model for Lean 4 proof assistant tasks. The model uses 119B total parameters with 6.5B activated per token, a 256k context length, and accepts text and image input. The source gives setup paths through Mistral Vibe and a local vLLM server, with recommended settings of temperature 1.0 and reasoning effort set to high for complex prompts. The model is licensed under Apache 2.0.

Jun 30

Jun 30Tue
  1. Nano Banana 2.1OfficialAI score38

    Google launches Nano Banana 2 Lite, its fastest image model yet

    AIGoogle's Nano Banana account announces Nano Banana 2 Lite, which it calls the fastest banana. According to the quoted Google DeepMind post, it generates text-to-image outputs in about 4 seconds and targets quicker ideation and workflows where speed and cost matter most.

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceOfficialAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    AIQwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 24

Jun 24Wed
  1. PaddlePaddleOfficialAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    AIPaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

    Image from @PaddlePaddle's post

Jun 19

Jun 19Fri

Jun 18

Jun 18Thu
  1. Cohere · new models on Hugging FaceOfficialAI score43

    Cohere Releases Open-Source 2B Arabic Speech Recognition Model Transcribe Arabic

    AICohere and Cohere Labs released Cohere Transcribe Arabic, an open-source 2B-parameter Arabic automatic speech recognition model under Apache 2.0. It is optimized for Arabic, Arabic dialects, English, and Arabic-English code-switched speech, using a Conformer encoder-decoder architecture supported natively in Transformers. The model's average WER of 25.87 and CER of 11.80 on the Open Universal Arabic ASR Leaderboard, as of 07.07.2026, is reported in the source.

Jun 16

Jun 16Tue
  1. Arthur MenschXAI score22

    Mistral plans a new sparse model family, with early access in July

    AIMistral says it will release a new model this summer, the first in a new family that is large but sparse. The company plans to open an early access program in July for key partners in research, government, and industry.

  2. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. Z.ai Release NotesOfficialAI score62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    AIZ.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

  2. ByteDance · new models on Hugging FaceOfficialAI score24

    Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

    AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 12

Jun 12Fri
  1. PaddlePaddleOfficialAI score41

    PaddleOCR releases PP-OCRv6 with models from 1.5M to 34.5M parameters

    AIPaddlePaddle has released PP-OCRv6, a new OCR model series in Tiny, Small, and Medium sizes at 1.5M, 7.7M, and 34.5M parameters. The models reportedly improve detection accuracy by 4.9% and recognition accuracy by 5.1% over PP-OCRv5, with up to 5.2× faster CPU inference via OpenVINO. The unified model supports 50 languages and new scenarios including PCB, CAD drawings, digital tubes, and dot-matrix text, under Apache 2.0.

    Image from @PaddlePaddle's post

Jun 11

Jun 11Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceOfficialAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

  2. ByteDance · new models on Hugging FaceOfficialAI score34

    EvoQuality: ByteDance's self-evolving VLM for image quality assessment without human labels

    AIEvoQuality is a ByteDance vision-language model for no-reference image quality assessment that generates pseudo-ranking labels through pairwise majority voting and refines them with GRPO, requiring no human-annotated quality scores. On the paper's setting, it raised weighted-average PLCC from 0.615 to 0.770 and SRCC from 0.570 to 0.726 over its Qwen2.5-VL-7B backbone. The model is recommended for research and pre-production assessment, not as the sole criterion for high-stakes decisions.

Jun 9

Jun 9Tue
  1. ByteDance · new models on Hugging FaceOfficialAI score28

    ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation

    AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.

  2. Andrej KarpathyXAI score65

    Karpathy Calls Claude Fable 5 a Major Step Forward for Long Tasks

    AIAndrej Karpathy says Claude Fable 5 is the same underlying model as Mythos with added safeguards, and that it leads on nearly all benchmarks. He describes it as a step change, especially for long, difficult problem-solving sessions where it handles more ambitious tasks without close supervision. He notes that its safeguards are set a bit too aggressively at launch and may be tuned over time.

  3. Z.ai (GLM) · new models on Hugging FaceOfficialAI score52

    Z.ai releases SCAIL-2, an open-source end-to-end character animation model

    AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.

Jun 8

Jun 8Mon
  1. Xiaomi MiMoOfficialAI score62

    Xiaomi MiMo-V2.5-Pro UltraSpeed claims 1,000+ tokens/s on a 1T model

    AIXiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, which the post says reaches output speeds above 1,000 tokens/s on a 1 trillion parameter MoE model. The post says this runs on a single standard 8-GPGPU node rather than wafer-scale or pure on-chip SRAM hardware. UltraSpeed access is application-based from Jun 8 to Jun 23 (PDT), and the UltraSpeed API costs 3x the standard price.

    Why it matters: The post specifies the hardware setup behind the claimed speed, which matters for judging whether the approach can be replicated on standard GPU nodes.

    Image from @XiaomiMiMo's post
  2. ByteDance · new models on Hugging FaceOfficialAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

  3. Xiaomi MiMo · new models on Hugging FaceOfficialAI score41

    Xiaomi releases MiMo-V2.5-Pro-FP4-DFlash, an FP4 model with block-diffusion decoding

    AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.

  4. Xiaomi MiMoOfficialAI score65

    Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model

    AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.

    Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceOfficialAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

  3. ByteDance · new models on Hugging FaceOfficialAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

May 31

May 31Sun
  1. MiniMax BlogOfficialAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat