Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 2

Sep 2Wed
  1. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  2. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  3. Unsloth AIOfficialAI score62

    Qwen3.8-Flash-Next runs 1.3 to 1.7 times faster locally with MTP

    AIUnsloth says MTP enables Qwen3.8-Flash-Next to run about 1.3 to 1.7 times faster at inference with no accuracy change. GGUF versions can reach 170 tokens/s on an RTX PRO 6000, and the source lists memory requirements from 76 GB at 1-bit to 355 GB at BF16.

    Why it matters: The source gives concrete MTP speedup ranges, hardware memory requirements, and GGUF quantization sizes, which help readers judge whether local deployment fits their setup.

    Image from @UnslothAI's post

Sep 1

Sep 1Tue
  1. Ai2 · new models on Hugging FaceOfficialAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  2. Google · new models on Hugging FaceOfficialAI score44

    Google Releases GNM v3.0, an Open 3D Parametric Model of the Human Head

    AIGoogle has released GNM v3.0, a parametric 3D statistical model of the human head, with weights published on Hugging Face and Kaggle under the Apache 2.0 license. The model gives controllable identity, expression, head pose, and internal anatomy including eyeballs, teeth, and tongue, and supports NumPy, JAX, PyTorch, and TensorFlow backends.

  3. Tencent HyOfficialAI score58

    Tencent Hy4 preview reports 31.8% throughput gain from self-found bottlenecks

    AITencent Hunyuan says its Hy4 preview model found inference bottlenecks on its own and raised end-to-end throughput by 31.8% through operator fusion and communication optimizations. The post says the gain holds across context lengths and concurrency levels. The release is listed at 770B total parameters with 49B active and a 1M context window, with links to the Hy blog, Hugging Face, and GitHub.

    Image from @TencentHunyuan's post
  4. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. The Register · AINewsAI score55

    OpenClaw 2.0 simplifies setup and adds shared sessions, but security defaults remain weak

    AIOpenClaw 2.0 is an open-source, self-hosted AI agent harness whose update simplifies installation, rebuilds the browser interface, and adds shared cloud sessions for multiple users. The article says the patch notes state shared session controls are not a security boundary, secret store values are not encrypted at rest, and sandboxing is off by default.

  2. Liquid AI NewsletterOfficialAI score46

    Liquid AI launches Pipette, an open-source benchmark for on-device foundation models

    AILiquid AI and Artificial Analysis released Pipette, an open-source benchmark platform for foundation models on edge devices, covering over 1,000 configurations across 30+ models. It measures five on-device metrics, including throughput, latency, context scaling, and memory use, on macOS, Windows, iOS, and Android. Liquid AI also said its updated LFM2.5 Q4_0 checkpoints, trained with Quantization-Aware Distillation, retain roughly 97% of BF16 baseline performance and suffer 73.4% less quality loss than standard post-training Q4_0 quantization.

  3. Microsoft ResearchOfficialAI score45

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Population-Scale Research

    AIMicrosoft Research released GigaPath-Flash and GigaTIME-Flash, efficient pathology foundation models built on a distilled ViT-S backbone and released under the Apache 2.0 license. GigaPath-Flash, with 22M-parameter tile and 21M-parameter slide encoders, reportedly scores within 3% of the original GigaPath on PANDA and EBRAINS benchmarks at roughly 50 times less compute. The models are research tools, not validated for clinical use.

  4. Amazon ScienceOfficialAI score45

    Amazon details using Verus to formally verify Rust code correctness

    AIAmazon Science explains Verus, an open-source automated program verifier for Rust that checks code against formal specifications for all possible inputs. Developers write specifications and proofs directly in Rust source using Rust-like syntax, and Verus returns feedback in under a second. Amazon says it has used Verus to prove the correctness of key primitives in the Nitro Isolation Engine and other infrastructure.

  5. DeepSeek · new models on Hugging FaceOfficialAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 29

Aug 29Sat
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score34

    Fun-ASR-Nano-2512 Gets vLLM-Native Packaging for Speech Transcription

    AIFunAudioLLM has released Fun-ASR-Nano-2512-vllm, a vLLM-native packaging of the official Fun-ASR-Nano-2512 checkpoint, with weights bitwise equal to the source and no new LoRA weights. The validated path runs on vLLM 0.27.1 with float32 through an OpenAI-compatible transcription endpoint, tested on one NVIDIA H100 80 GB GPU. The source-licensed model is Apache License 2.0, and other vLLM versions, accelerators, and quantizations require separate validation.

  2. Tencent HyOfficialAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri
  1. TinkerOfficialAI score52

    GLM-5.3 from Z.ai is now available on Tinker with 256k context

    AITinker announces that Z.ai's GLM-5.3 is now available on its platform with a 256k context window. Tinker says it is currently the strongest open-weights model on coding evals including Terminal-Bench 3.0 and DeepSWE 1.1, built on the same base as GLM-5.2 with scaled-up post-training.

  2. Daniel HanXAI score50

    Unsloth quantizes GLM-5.3 to 1-bit at 217GB with 76% accuracy retained

    AIUnsloth quantized GLM-5.3 to a dynamic 1-bit version of 217GB, versus 1.5TB for BF16, retaining about 76% top-1 accuracy while cutting size by 83%. The team said a 1-bit build ran a simple snake game well in Unsloth Desktop. The post credits Z.ai's GLM-5.3-Flash and GLM-5.3 releases.

    Video from @danielhanchen's post
  3. Unsloth AIOfficialAI score70

    Unsloth shows how to run GLM-5.3 locally with 2-bit quantization

    AIUnsloth AI published a guide for running GLM-5.3 locally using quantized GGUF weights. The 2-bit version is reduced from 1.51TB to 239GB and retains about 81% accuracy, and it can run on a 256GB Mac or RAM/VRAM setups.

    Why it matters: The guide shows which quantization levels fit local memory budgets and how much accuracy each costs, useful for planning a local deployment.

    Image from @UnslothAI's post
  4. LM StudioOfficialAI score38

    GLM-5.3 Now Available in Bionic Agent, 50% Off Until Monday

    AILM Studio says Zhipu's GLM-5.3 is now available in its Bionic agent, hosted in the US with zero data retention. The launch offer is 50% off until Monday. The post builds on Z.ai's announcement that GLM-5.3 is now open-weight for agentic coding and cyber defense.

  5. LMSYS OrgOfficialAI score42

    SGLang adds day-0 serving support for GLM-5.3 across NVIDIA and AMD GPUs

    AISGLang offers day-0 support for GLM-5.3 on NVIDIA Blackwell and Hopper and AMD MI300X, MI325X, and MI355X GPUs, using the same runtime and flags. SGLang is also the rollout engine in Slime, the framework Zhipu used to post-train GLM-5.3, so the runtime that generated the RL trajectories now serves the model.

  6. LMSYS OrgOfficialAI score34

    Infer-forge: Three-layer agent system for SGLang inference optimization

    AIAnt OSS built Infer-forge, a three-layer system of Harness, Task Loop, and Task Graph that runs long SGLang inference optimization work through agents while keeping provenance. Peak Tasks in flight rose from 2 to 9, and median Task lifetime grew from 10 hours to 28 hours. The agent independently ran a full serving project on DeepSeek-V4-Pro, splitting the work into 38 verified pieces and catching kernel silent corruption on its own.

    Image from @lmsysorg's post
  7. RadixArkOfficialAI score40

    RadixArk releases experimental NVFP4 checkpoint for GLM-5.3

    AIRadixArk has published an experimental NVFP4 checkpoint for Zhipu's GLM-5.3 on Hugging Face, and says Miles support for GLM-5.3 is on the way. The company says GLM-5.2 is already serving hundreds of thousands of people in production with its partners on SGLang. Background from SGLang reports day-0 serving support for GLM-5.3, with 537.6 tok/s/user on NVFP4 and 413 tok/s/user on FP8 at BS=1 with TP8 on 8x B300.

  8. Z.aiOfficialAI score62

    Z.ai releases GLM-5.3 as open-weight model for agentic coding and cyber defense

    AIZ.ai has made GLM-5.3 open-weight, so users can download, run, and customize its weights. The company describes it as its most capable model for agentic coding and cyber defense, with weights on Hugging Face and details in a tech blog.

    Why it matters: The source ties the open-weight release to coding, agentic, and cyber defense use, with a weights link for anyone who wants to run or customize it.

    Image from @Zai_org's post

Aug 27

Aug 27Thu
  1. LMSYS OrgOfficialAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

    Image from @lmsysorg's post
  2. Daniel HanXAI score36

    GLM-5.3-Flash quantizes to 4-bit with 93% accuracy retained

    AIGLM-5.3-Flash (ox-alpha) can be quantized to 4-bit while retaining 93% accuracy, according to Daniel Han. The post says the 4-bit model runs on a 256GB Mac or two DGX Sparks, and 5-bit may also work. Unsloth separately says 3-bit GGUF runs on 128GB RAM and that the model rivals Claude Opus 4.8 on DeepSWE, coding, and agentic benchmarks.

    Image from @danielhanchen's post
  3. Unsloth AIOfficialAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  4. Leandro von WerraXAI score22

    Pollen Robotics unveils Microduck, a $400 open-source RL biped robot

    AIPollen Robotics has unveiled Microduck, a 25 cm open-source biped with 15 actuators and sensors including a camera, speaker, and LiDAR that users can train with reinforcement learning. The robot ships with more than half a dozen pre-trained policies for walking, sitting, roller-skating, and picking up objects with its articulated beak, and costs less than $400.

  5. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  6. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

  7. Tencent · new models on Hugging FaceOfficialAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    AITencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    Why it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  8. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceOfficialAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.