Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 29

Aug 29Sat
  1. Tencent HyAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri

Aug 27

Aug 27Thu
  1. Unsloth AIAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  3. OpenBMB (MiniCPM) · new models on Hugging FaceAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

  4. Tencent · new models on Hugging FaceAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    AITencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    Why it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  5. Qwen · new models on Hugging FaceAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    AITencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  3. LMSYS OrgAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  4. Unsloth AIAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    AIUnsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    Why it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

    Image from @UnslothAI's post
  5. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  6. Ai2 · new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  7. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  8. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Z.ai Release NotesAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  3. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  4. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging FaceAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    AIGoogle Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Microsoft ResearchAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post
  3. Qwen · new models on Hugging FaceAI score75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    AIQwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.