Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 11

Aug 11Tue
  1. Liquid AI · new models on Hugging FaceOfficialAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

  2. koray kavukcuogluXAI score38

    Gemini app surpasses 1 billion monthly users, Google says

    AIGoogle's Gemini app has reached over 1 billion monthly users, according to Koray Kavukcuoglu's post celebrating the milestone. Sundar Pichai said it is Google's fastest-growing product ever and the company's 14th product to reach 1 billion users.

  3. Rowan CheungXAI score62

    Meta opens weights for Muse Glimmer 30B model, Muse Spark 1.2 to follow

    AIMeta announced it is opening the weights for Muse Glimmer, a 30B parameter dense model that can run locally. Muse Spark 1.2, described as its latest foundation model, will have its weights released soon. The author's interview with Mark Zuckerberg quotes him saying Llama 4 fell short of the trajectory he wanted and that the lab was rebuilt.

    Video from @rowancheung's post
  4. Bryan CatanzaroXAI score40

    Nemotron 3.5 Lightning: NVIDIA's fast 30B MoE model for agents

    AINVIDIA's Bryan Catanzaro says Nemotron 3.5 Lightning uses the same architecture as Nemotron 3.0 Nano, adds speculative decoding, and matches the intelligence of Nemotron 3.0 Super. NVIDIA describes it as an open 30B MoE model with 3B active parameters, built for always-on agents handling high-volume, specialized tasks, with up to 4x the output speed of similar-sized models.

Aug 10

Aug 10Mon
  1. Liquid AI · new models on Hugging FaceOfficialAI score43

    LiquidAI LFM2.5-2.6B-DSpark Speeds Up LFM2.5 Decoding With Speculative Drafting

    AILiquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face. In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x. Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.

  2. Liquid AI · new models on Hugging FaceOfficialAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  3. Liquid AI · new models on Hugging FaceOfficialAI score42

    LiquidAI LFM2.5-1.2B-Instruct-DSpark Drafter Speeds Up Decoding About 2x

    AILiquid AI released LFM2.5-1.2B-Instruct-DSpark, a 295.7M-parameter speculative-decoding draft model for the LFM2.5-1.2B-Instruct target on Hugging Face. On an H100 it averages 4.81 accepted tokens per step and runs about 2.10x faster across benchmarks, with about 2x speedup in SGLang and on-device Apple silicon support via Metal.

  4. Cohere · new models on Hugging FaceOfficialAI score46

    Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model

    AICohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face. The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.

Aug 9

Aug 9Sun
  1. Fireworks AI BlogOfficialAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceOfficialAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. MiniMax · new models on Hugging FaceOfficialAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

Aug 6

Aug 6Thu
  1. Intern Large ModelsOfficialAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

    Why it matters: The post links Mobius to the RNN-to-Transformer shift and gives speed, data, and generalization figures that readers can weigh against standard Transformers.

    Image from @intern_lm's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceOfficialAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

  2. Meituan LongCatOfficialAI score34

    LongCat-2.0 is now free on OpenCode for coding

    AIMeituan's LongCat-2.0, a coding-optimized model with 1M context and full open-source availability, is now free to use on OpenCode. The announcement comes from Meituan LongCat's official X account and does not include pricing details beyond the free offer.

Aug 3

Aug 3Mon
  1. Liquid AI BlogOfficialAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

Jul 31

Jul 31Fri
  1. Thinking MachinesOfficialAI score44

    Thinking Machines argues for staged access to capable open-weight models

    AIThinking Machines says indiscriminately releasing model weights is unsafe, but keeping capable models inside a few labs is also not the answer. Its new post describes how it assessed its model Inkling and argues that access should widen in stages. The company says it has not mapped the full path, only the portion it can currently see.

  2. DeepSeek · new models on Hugging FaceOfficialAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  3. DeepSeekOfficialAI score38

    DeepSeek-V4-Flash-0731 API upgrade keeps preview architecture and size

    AIDeepSeek says DeepSeek-V4-Flash-0731 keeps the same model architecture and size as the preview version. Today's upgrade applies only to the DeepSeek-V4-Flash API, while the DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official DeepSeek-V4-Pro release is coming soon.

  4. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogOfficialAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Thinking Machines LabOfficialAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    AIThinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    Why it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

  3. Soumith ChintalaXAI score57

    Thinking Machines releases Inkling-Small, a 276B-parameter model with full weights

    AIThinking Machines is releasing Inkling-Small, which it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and the full weights are available. Users can fine-tune it on Tinker or chat with it in text, image, and audio on Tinker Playground.

  4. Mira MuratiXAI score62

    Thinking Machines releases open-weight Inkling-Small model at a quarter of Inkling's size

    AIThinking Machines has released Inkling-Small, which the author says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and its full weights are available, with fine-tuning on Tinker and chat in text, image, and audio on Tinker Playground.

    Why it matters: The post states Inkling-Small's size and open-weight status, helping readers compare it with the larger Inkling and judge fine-tuning access.

  5. Thinking MachinesOfficialAI score43

    Thinking Machines' new model matches Inkling on multimodal evals

    AIThinking Machines' new model is natively multimodal and encoder-free, processing audio and images jointly with text. It nearly matches Inkling across multimodal evaluations and can run Python to crop, zoom, and inspect images while reasoning over documents and charts.

  6. Thinking MachinesOfficialAI score44

    Thinking Machines' Inkling-Small Gains From Lessons of Larger Inkling

    AIThinking Machines says Inkling-Small began training after its larger counterpart and benefits from the lessons learned. The post cites an improved pre-training data mix, a refined ML recipe, on-policy distillation using Inkling as the teacher, and two additional weeks of agentic coding RL.

  7. Thinking MachinesOfficialAI score62

    Thinking Machines releases Inkling-Small, with full weights available

    AIThinking Machines is releasing Inkling-Small, a model it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and full weights are available. It can be fine-tuned on Tinker or used for text, image, and audio chat in the Tinker Playground.

    Why it matters: The release gives comparable performance at a quarter of Inkling's size, with full weights available, which matters for teams weighing self-hosting against larger models.

  8. Thinking MachinesOfficialAI score38

    Thinking Machines' Inkling-Small gains performance per FLOP over Inkling

    AIThinking Machines says its Inkling-Small model delivers more performance per FLOP than Inkling on Terminal-Bench 2.1 agentic tool use, HLE reasoning, and IFBench instruction following. Variable thinking effort lets users choose their own point on the cost-performance curve.

  9. IdeogramOfficialAI score43

    Ideogram and Pruna launch P-Image-Ideogram image model family

    AIIdeogram introduced P-Image-Ideogram, a family of image models co-developed with Pruna, offering a quality-speed-cost trade-off across four quality modes. The models generate native 1K and 2K images starting at $0.003 per image and are available now on the Ideogram API and partner platforms.

    Video from @ideogram_ai's post

Jul 29

Jul 29Wed
  1. Google LabsOfficialAI score44

    Google DeepMind releases Lyria 3.5, integrated into Flow Music

    AIGoogle DeepMind has released Lyria 3.5 and integrated it directly into Flow Music. The update improves prompt adherence so users can set exact BPMs and export stems for full-length songs, along with more expressive vocals and more natural note-to-note musical arrangements.

    Video from @GoogleLabs's post
  2. Google LabsOfficialAI score60

    Google Launches Lyria 3.5 in Flow Music With Better Vocals and Lyrics

    AIGoogle is rolling out Lyria 3.5, its newest music generation model, in Google Flow Music today. The update improves musicality, lyric quality and prompt adherence, and vocal expressiveness and pronunciation, and gives users more control over tempo and duration.

    Why it matters: The post names the specific capability changes and where users can access them, which helps readers judge fit for music creation workflows.

  3. Air Street PressBlogAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

    Why it matters: The piece ties Laguna S 2.1's open weights and published trajectories to Poolside's release cadence, showing how its model factory compounds gains across successive releases.

  4. Liquid AI NewsletterOfficialAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  6. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  7. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Augment Code BlogOfficialAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.