Skip to contentSkip to stories

Updated

#Model release

Aug 11

Aug 11Tue
  1. Liquid AI BlogAI score62

    Liquid AI releases LFM2.5-VL-3B, a 3B vision-language model for edge devices

    AILiquid AI released LFM2.5-VL-3B, an open-weight 3B vision-language model that it says rivals models twice its size while running faster on CPU and GPU. Benchmarks show large gains over LFM2-VL-3B, including ScreenSpot-v2 averaging 80.7, RefCOCO precision@1 rising from 57.1 to 87.9, and ToolSandbox rising from 26.4 to 59.5. The model is available on Hugging Face and decodes 228 tokens/s on an Apple M5 Max.

    Why it matters: The post pairs benchmark gains with on-device and GPU throughput figures, showing how a 3B vision model trades size against speed and accuracy.

  2. Liquid AI · new models on Hugging FaceAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

  3. Rowan CheungAI score62

    Meta opens weights for Muse Glimmer 30B model, Muse Spark 1.2 to follow

    AIMeta announced it is opening the weights for Muse Glimmer, a 30B parameter dense model that can run locally. Muse Spark 1.2, described as its latest foundation model, will have its weights released soon. The author's interview with Mark Zuckerberg quotes him saying Llama 4 fell short of the trajectory he wanted and that the lab was rebuilt.

  4. Bryan CatanzaroAI score40

    Nemotron 3.5 Lightning: NVIDIA's fast 30B MoE model for agents

    AINVIDIA's Bryan Catanzaro says Nemotron 3.5 Lightning uses the same architecture as Nemotron 3.0 Nano, adds speculative decoding, and matches the intelligence of Nemotron 3.0 Super. NVIDIA describes it as an open 30B MoE model with 3B active parameters, built for always-on agents handling high-volume, specialized tasks, with up to 4x the output speed of similar-sized models.

Aug 10

Aug 10Mon
  1. Liquid AI · new models on Hugging FaceAI score43

    LiquidAI LFM2.5-2.6B-DSpark Speeds Up LFM2.5 Decoding With Speculative Drafting

    AILiquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face. In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x. Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.

  2. Liquid AI · new models on Hugging FaceAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  3. Liquid AI · new models on Hugging FaceAI score42

    LiquidAI LFM2.5-1.2B-Instruct-DSpark Drafter Speeds Up Decoding About 2x

    AILiquid AI released LFM2.5-1.2B-Instruct-DSpark, a 295.7M-parameter speculative-decoding draft model for the LFM2.5-1.2B-Instruct target on Hugging Face. On an H100 it averages 4.81 accepted tokens per step and runs about 2.10x faster across benchmarks, with about 2x speedup in SGLang and on-device Apple silicon support via Metal.

  4. Cohere · new models on Hugging FaceAI score46

    Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model

    AICohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face. The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.

Aug 9

Aug 9Sun
  1. Fireworks AI BlogAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. MiniMax · new models on Hugging FaceAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

Aug 6

Aug 6Thu
  1. Intern Large ModelsAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

Aug 3

Aug 3Mon
  1. Liquid AI BlogAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

Aug 1

Aug 1Sat

Jul 31

Jul 31Fri
  1. Thinking MachinesAI score44

    Thinking Machines argues for staged access to capable open-weight models

    AIThinking Machines says indiscriminately releasing model weights is unsafe, but keeping capable models inside a few labs is also not the answer. Its new post describes how it assessed its model Inkling and argues that access should widen in stages. The company says it has not mapped the full path, only the portion it can currently see.

  2. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  3. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Thinking Machines LabAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    AIThinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    Why it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

Jul 29

Jul 29Wed
  1. Google LabsAI score60

    Google Launches Lyria 3.5 in Flow Music With Better Vocals and Lyrics

    AIGoogle is rolling out Lyria 3.5, its newest music generation model, in Google Flow Music today. The update improves musicality, lyric quality and prompt adherence, and vocal expressiveness and pronunciation, and gives users more control over tempo and duration.

    Why it matters: The post names the specific capability changes and where users can access them, which helps readers judge fit for music creation workflows.

  2. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  3. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  6. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.