Skip to content

#Open source/Repo

Sep 1

Sep 1Tue
  1. Tencent HunyuanAI score58

    Tencent Hy4 preview reports 31.8% throughput gain from self-found bottlenecks

    Tencent Hunyuan says its Hy4 preview model found inference bottlenecks on its own and raised end-to-end throughput by 31.8% through operator fusion and communication optimizations. The post says the gain holds across context lengths and concurrency levels. The release is listed at 770B total parameters with 49B active and a 1M context window, with links to the Hy blog, Hugging Face, and GitHub.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    Shanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    AIWhy it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. The Register · AIAI score55

    OpenClaw 2.0 simplifies setup and adds shared sessions, but security defaults remain weak

    OpenClaw 2.0 is an open-source, self-hosted AI agent harness whose update simplifies installation, rebuilds the browser interface, and adds shared cloud sessions for multiple users. The article says the patch notes state shared session controls are not a security boundary, secret store values are not encrypted at rest, and sandboxing is off by default.

  2. Amazon ScienceAI score45

    Amazon details using Verus to formally verify Rust code correctness

    Amazon Science explains Verus, an open-source automated program verifier for Rust that checks code against formal specifications for all possible inputs. Developers write specifications and proofs directly in Rust source using Rust-like syntax, and Verus returns feedback in under a second. Amazon says it has used Verus to prove the correctness of key primitives in the Nitro Isolation Engine and other infrastructure.

  3. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    DeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    AIWhy it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    Alibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    Alibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

Aug 29

Aug 29Sat
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score34

    Fun-ASR-Nano-2512 Gets vLLM-Native Packaging for Speech Transcription

    FunAudioLLM has released Fun-ASR-Nano-2512-vllm, a vLLM-native packaging of the official Fun-ASR-Nano-2512 checkpoint, with weights bitwise equal to the source and no new LoRA weights. The validated path runs on vLLM 0.27.1 with float32 through an OpenAI-compatible transcription endpoint, tested on one NVIDIA H100 80 GB GPU. The source-licensed model is Apache License 2.0, and other vLLM versions, accelerators, and quantizations require separate validation.

Aug 28

Aug 28Fri
  1. Daniel HanAI score50

    We quantized GLM-5.3 to dynamic 1-bit (217GB) vs BF16 (1.5TB) and it retains ~76% top-1% accuracy whilst being 83% We did a small basic snake game in Unsloth Desktop using the UD 1-bit and it worked well! zai is on a roll with GLM-5.3-Flash and now GLM-5.3!

    We quantized GLM-5.3 to dynamic 1-bit (217GB) vs BF16 (1.5TB) and it retains ~76% top-1% accuracy whilst being 83% We did a small basic snake game in Unsloth Desktop using the UD 1-bit and it worked well! zai is on a roll with GLM-5.3-Flash and now GLM-5.3!

  2. Unsloth AIAI score70

    Unsloth shows how to run GLM-5.3 locally with 2-bit quantization

    Unsloth AI published a guide for running GLM-5.3 locally using quantized GGUF weights. The 2-bit version is reduced from 1.51TB to 239GB and retains about 81% accuracy, and it can run on a 256GB Mac or RAM/VRAM setups.

    AIWhy it matters: The guide shows which quantization levels fit local memory budgets and how much accuracy each costs, useful for planning a local deployment.

Aug 27

Aug 27Thu
  1. LMSYS OrgAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    MiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

  2. VercelAI score40

    We built https://vgpu.sh to ship shaders on https://vercel.com. Now it's open source. ▪︎ Minimal agent-first WebGPU library ▪︎ Run in the browser or headless Node.js ▪︎ Render in CPU sandboxes and CI tests ▪︎ Create reusable .wgsl modules

    We built https://vgpu.sh to ship shaders on https://vercel.com. Now it's open source. ▪︎ Minimal agent-first WebGPU library ▪︎ Run in the browser or headless Node.js ▪︎ Render in CPU sandboxes and CI tests ▪︎ Create reusable .wgsl modules

  3. Daniel HanAI score36

    GLM-5.3-Flash (ox-alpha) can be quantized down to 4-bit and retain 93% accuracy! The 4-bit model is essentially like a local Claude 4.7 Opus. It runs perfectly on a 256GB Mac or two DGX Sparks. 5-bit may even work.

    GLM-5.3-Flash (ox-alpha) can be quantized down to 4-bit and retain 93% accuracy! The 4-bit model is essentially like a local Claude 4.7 Opus. It runs perfectly on a 256GB Mac or two DGX Sparks. 5-bit may even work.

  4. Unsloth AIAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    Unsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    AIWhy it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

  5. Tencent · new models on Hugging FaceAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    Tencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    AIWhy it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  6. Qwen · new models on Hugging FaceAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    Qwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    AIWhy it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    Tencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    Tencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  3. LMSYS OrgAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    Z.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    AIWhy it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  4. Unsloth AIAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    Unsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    AIWhy it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

  5. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    Ai2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  6. Ai2 · new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    Ai2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  7. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    Ai2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Andrew NgAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    OpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  2. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    Z.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    AIWhy it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging FaceAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    Google Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Thinking MachinesAI score34

    Today, we are launching Tinker grants of up to $50,000 in credits for safety research on open-weight models. We share some project ideas that excite us below; if you’re working on a safety project that could be accelerated by additional Tinker credits, we want to hear from you!

    Today, we are launching Tinker grants of up to $50,000 in credits for safety research on open-weight models. We share some project ideas that excite us below; if you’re working on a safety project that could be accelerated by additional Tinker credits, we want to hear from you!

  3. Engineering at MetaAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    Meta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    AIWhy it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.