Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Jul 31

Jul 31Fri
  1. Thinking MachinesOfficialAI score44

    Thinking Machines argues for staged access to capable open-weight models

    AIThinking Machines says indiscriminately releasing model weights is unsafe, but keeping capable models inside a few labs is also not the answer. Its new post describes how it assessed its model Inkling and argues that access should widen in stages. The company says it has not mapped the full path, only the portion it can currently see.

  2. DeepSeek · new models on Hugging FaceOfficialAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  3. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogOfficialAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Thinking Machines LabOfficialAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    AIThinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    Why it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

  3. Thinking MachinesOfficialAI score43

    Thinking Machines' new model matches Inkling on multimodal evals

    AIThinking Machines' new model is natively multimodal and encoder-free, processing audio and images jointly with text. It nearly matches Inkling across multimodal evaluations and can run Python to crop, zoom, and inspect images while reasoning over documents and charts.

  4. Thinking MachinesOfficialAI score44

    Thinking Machines' Inkling-Small Gains From Lessons of Larger Inkling

    AIThinking Machines says Inkling-Small began training after its larger counterpart and benefits from the lessons learned. The post cites an improved pre-training data mix, a refined ML recipe, on-policy distillation using Inkling as the teacher, and two additional weeks of agentic coding RL.

  5. Thinking MachinesOfficialAI score62

    Thinking Machines releases Inkling-Small, with full weights available

    AIThinking Machines is releasing Inkling-Small, a model it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and full weights are available. It can be fine-tuned on Tinker or used for text, image, and audio chat in the Tinker Playground.

  6. IdeogramOfficialAI score43

    Ideogram and Pruna launch P-Image-Ideogram image model family

    AIIdeogram introduced P-Image-Ideogram, a family of image models co-developed with Pruna, offering a quality-speed-cost trade-off across four quality modes. The models generate native 1K and 2K images starting at $0.003 per image and are available now on the Ideogram API and partner platforms.

    Video from @ideogram_ai's post

Jul 29

Jul 29Wed
  1. Google LabsOfficialAI score44

    Google DeepMind releases Lyria 3.5, integrated into Flow Music

    AIGoogle DeepMind has released Lyria 3.5 and integrated it directly into Flow Music. The update improves prompt adherence so users can set exact BPMs and export stems for full-length songs, along with more expressive vocals and more natural note-to-note musical arrangements.

    Video from @GoogleLabs's post
  2. Google LabsOfficialAI score60

    Google Launches Lyria 3.5 in Flow Music With Better Vocals and Lyrics

    AIGoogle is rolling out Lyria 3.5, its newest music generation model, in Google Flow Music today. The update improves musicality, lyric quality and prompt adherence, and vocal expressiveness and pronunciation, and gives users more control over tempo and duration.

    Why it matters: The post names the specific capability changes and where users can access them, which helps readers judge fit for music creation workflows.

  3. Air Street PressBlogAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  4. Liquid AI NewsletterOfficialAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  6. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  7. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Augment Code BlogOfficialAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. MiniMax · new models on Hugging FaceOfficialAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

  3. Intern Large ModelsOfficialAI score62

    Intern Large Models introduces Visual Pretraining learned from visual documents

    AIIntern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents. The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

    Image from @intern_lm's post

Jul 27

Jul 27Mon
  1. Liquid AI BlogOfficialAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    AILiquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  2. Kimi.aiOfficialAI score65

    Kimi K3 becomes available on Nebius Token Factory via API

    AIKimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.

    Why it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.

    Image from @Kimi_Moonshot's post
  3. Kimi.aiOfficialAI score35

    Kimi K3 launches day 0 on Baseten's Model APIs

    AIMoonshot AI's Kimi K3 is available from day zero through Baseten's Model APIs, offering fast and reliable access. Baseten is named as the launch partner for bringing K3 to more users.

    Image from @Kimi_Moonshot's post
  4. Kimi.aiOfficialAI score47

    Kimi K3 launches with Modal as Day 0 partner for faster inference

    AIKimi K3 is available on Modal as a Day 0 launch partner, with Modal training a custom DFlash speculator for the model's architecture. The speculator delivers faster inference with no quality loss, according to Kimi. Modal describes K3 as a 3T-class open model that is the most capable open model it has worked with.

    Image from @Kimi_Moonshot's post
  5. Kimi.aiOfficialAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

    Image from @Kimi_Moonshot's post
  6. Air Street PressBlogAI score72

    Black Forest Labs releases FLUX 3, extended to video and robot control

    AIBlack Forest Labs released FLUX 3, a multimodal model trained on images, video, and audio, and mimic built FLUX-mimic on its video backbone to control robots. In a soft-body kitting task, mimic reports a 95% success rate without single-task fine-tuning, compared with 55% for an adapted π0.5 model. FLUX 3 Video is in early access, with action prediction offered to selected partners and an open-weight backbone planned.

Jul 26

Jul 26Sun
  1. Fireworks AI BlogOfficialAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

Jul 25

Jul 25Sat
  1. Sebastien BubeckXAI score25

    Bubeck asks what Erdős would do with GPT-5.6 Sol

    AISebastien Bubeck asked what the mathematician Paul Erdős would have done with access to GPT-5.6 Sol. The post offers no benchmark results, prices, or capability details beyond the question itself.

Jul 24

Jul 24Fri
  1. MidjourneyOfficialAI score60

    Midjourney releases V8.2 as its new default image model

    AIMidjourney is releasing V8.2 today and making it the default model on the platform. The release focuses on aesthetics, personalization, and image quality, with a bolder and more creative new style. The post's two image pairs compare earlier outputs with the new style.

    Image from @midjourney's post
  2. catXAI score66

    Claude Opus 5 released as strong option for long-running autonomous work

    AIAnthropic introduces Claude Opus 5 as a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price, according to the quoted announcement. The author, who works on the product, says Claude Opus 5 is great at long-running autonomous work and invites users to try it and share feedback.

    Why it matters: The post pairs a new model's long-running autonomous strength with a pricing claim, letting readers weigh capability against cost for agentic workloads.

  3. Alex AlbertXAI score72

    Anthropic introduces Claude Opus 5, close to Fable 5 intelligence at half the price

    AIAnthropic has introduced Claude Opus 5, which the quoted announcement describes as a thoughtful and proactive model. It is said to come close to the frontier intelligence of Fable 5 at half the price.

    Why it matters: The quoted announcement gives a concrete comparison of intelligence and price against Fable 5, useful for judging where Opus 5 fits among Claude models.

Jul 23

Jul 23Thu
  1. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

  2. BAAI · new models on Hugging FaceOfficialAI score47

    BAAI releases AREX-Turbo, a compact 4B recursive self-improving deep research agent

    AIBAAI's AREX-Turbo is a dense 4B deep research agent built on Qwen3.5-4B with a 262,144-token context length. It scores 70.7 on BrowseComp, 81.6 on GAIA and 40.6 on HLE with tools, versus 82.5, 85.4 and 52.4 for the 122B AREX-Base. The model is released under Apache License 2.0 and targets lower-cost research-agent deployment.

Jul 21

Jul 21Tue
  1. Bryan CatanzaroXAI score57

    Poolside releases open-weight Laguna S 2.1 for agentic coding

    AIPoolside released Laguna S 2.1, an open-weight model with 118B total parameters and 8B active per token. The author says it performs strongly on agentic coding and long-horizon tasks, and it can run on a single NVIDIA DGX Spark. Weights are on Hugging Face under the OpenMDW-1.1 license, with access also available through OpenRouter and Poolside's API.