Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 29

Aug 29Sat
  1. Tencent HyOfficialAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri
  1. TinkerOfficialAI score52

    GLM-5.3 from Z.ai is now available on Tinker with 256k context

    AITinker announces that Z.ai's GLM-5.3 is now available on its platform with a 256k context window. Tinker says it is currently the strongest open-weights model on coding evals including Terminal-Bench 3.0 and DeepSWE 1.1, built on the same base as GLM-5.2 with scaled-up post-training.

  2. Daniel HanXAI score50

    Unsloth quantizes GLM-5.3 to 1-bit at 217GB with 76% accuracy retained

    AIUnsloth quantized GLM-5.3 to a dynamic 1-bit version of 217GB, versus 1.5TB for BF16, retaining about 76% top-1 accuracy while cutting size by 83%. The team said a 1-bit build ran a simple snake game well in Unsloth Desktop. The post credits Z.ai's GLM-5.3-Flash and GLM-5.3 releases.

    Video from @danielhanchen's post
  3. Unsloth AIOfficialAI score70

    Unsloth shows how to run GLM-5.3 locally with 2-bit quantization

    AIUnsloth AI published a guide for running GLM-5.3 locally using quantized GGUF weights. The 2-bit version is reduced from 1.51TB to 239GB and retains about 81% accuracy, and it can run on a 256GB Mac or RAM/VRAM setups.

    Why it matters: The guide shows which quantization levels fit local memory budgets and how much accuracy each costs, useful for planning a local deployment.

    Image from @UnslothAI's post
  4. LM StudioOfficialAI score38

    GLM-5.3 Now Available in Bionic Agent, 50% Off Until Monday

    AILM Studio says Zhipu's GLM-5.3 is now available in its Bionic agent, hosted in the US with zero data retention. The launch offer is 50% off until Monday. The post builds on Z.ai's announcement that GLM-5.3 is now open-weight for agentic coding and cyber defense.

  5. MidjourneyOfficialAI score42

    Midjourney begins testing its first V8.2 edit model today

    AIMidjourney is starting tests of its first V8.2 edit model, which supports instruction-based editing and generating images from up to 4 reference images. The model also offers brush-based inpainting and outpainting, and works with personalization, moodboards, and srefs.

    Image from @midjourney's post

Aug 27

Aug 27Thu
  1. ReplicateOfficialAI score10

    Google's Gemini Omni 1.1 now available to try on Replicate

    AIReplicate has made Google's Gemini Omni 1.1 available to try through a hosted model page. The post provides only a link to the model on replicate.com, with no specifications, benchmarks, or pricing mentioned.

  2. Daniel HanXAI score36

    GLM-5.3-Flash quantizes to 4-bit with 93% accuracy retained

    AIGLM-5.3-Flash (ox-alpha) can be quantized to 4-bit while retaining 93% accuracy, according to Daniel Han. The post says the 4-bit model runs on a 256GB Mac or two DGX Sparks, and 5-bit may also work. Unsloth separately says 3-bit GGUF runs on 128GB RAM and that the model rivals Claude Opus 4.8 on DeepSWE, coding, and agentic benchmarks.

    Image from @danielhanchen's post
  3. Unsloth AIOfficialAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  4. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  5. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

  6. Tencent · new models on Hugging FaceOfficialAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    AITencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    Why it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  7. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Bryan CatanzaroXAI score46

    NVIDIA Releases DLSS 4.5 Ray Reconstruction with Better Image Quality

    AINVIDIA's DLSS 4.5 Ray Reconstruction is now available, using a second-generation joint denoiser and super-resolution model. According to the post, it delivers much better image quality at the same compute cost, pushing the trade-off between image quality and rendering cost further.

  2. Tencent · new models on Hugging FaceOfficialAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  3. Tencent · new models on Hugging FaceOfficialAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    AITencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  4. LM StudioOfficialAI score57

    GLM-5.3-Flash by Z.ai is now live in LM Studio

    AILM Studio announced that Z.ai's GLM-5.3-Flash, previously previewed as Ox Alpha, is available in LM Studio Bionic. The source says the model outperforms GLM-5.2 at 9-10x lower cost, supports image input, and is served from US-based servers with ZDR enabled by default.

  5. LM StudioOfficialAI score18

    LM Studio adds GLM-5.3 Flash model page

    AILM Studio published a model page for GLM-5.3 Flash on its website. The post contains only a link to the model page and no further details about the model's specifications or capabilities.

  6. Google AI DevelopersOfficialAI score30

    Google AI Developers points to Gemini 3.5 Transcribe resources

    AIThe post links to a Google blog page about Gemini 3.5 Transcribe, presented as a model developers can start building with. The post itself gives no benchmark scores, pricing, or capability details beyond the link.

  7. Google AI DevelopersOfficialAI score47

    Google launches Gemini 3.5 Transcribe, a speech-to-text model for developers

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that filters out spoken hesitations and accurately grounds technical terms, file names, and code variables against the active context. The model uses visual biasing to incorporate screen-aware context into developer workflows, as demonstrated in Antigravity.

    Video from @googleaidevs's post
  8. Sundar PichaiXAI score42

    Google launches Gemini 3.5 Transcribe with 85+ language support

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that auto-detects over 85 languages and handles multiple speakers. It also supports custom vocabulary adaptation for specialized jargon. The API is available now in Google AI Studio and Gemini Enterprise.

    Video from @sundarpichai's post
  9. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  10. Unsloth AIOfficialAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    AIUnsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    Why it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

    Image from @UnslothAI's post
  11. LMSYS OrgOfficialAI score60

    SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

    AISGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

  12. GeneralistOfficialAI score22

    GEN-1.5 shows physical prompt steerability in a shared environment

    AIGeneralist AI demonstrated that GEN-1.5 behaves differently when given different physical prompts in the same environment. The post shares a simple example responding to a request from @JagdeepBhatia8 about whether the model follows physical prompts or just the most likely action for the scene.

    Video from @GeneralistAI's post
  13. Z.aiOfficialAI score34

    GLM-5.3-Flash delivers greater intelligence with less compute

    AIZ.ai says GLM-5.3-Flash achieves greater intelligence with less compute, attributing this to architectural enhancements and an optimized pre-training corpus. The post gives no benchmark scores, parameter counts, or pricing.

    Image from @Zai_org's post
  14. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  15. Ai2 · new models on Hugging FaceOfficialAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  16. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  17. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

  18. Ai2 · new models on Hugging FaceOfficialAI score39

    Ai2 releases Bolmo-1B-Stage1, a byte-level version of OLMo 2 1B

    AIAi2 has released Bolmo-1B-Stage1, a 1B-parameter byte-level language model retrofitted from OLMo 2 1B to process bytes rather than tokens. This checkpoint includes Stage 1 training only, with inner model parameters unchanged, and is available on Hugging Face under an Apache 2.0 license for research and educational use.

  19. Ai2 · new models on Hugging FaceOfficialAI score40

    Ai2 Releases Bolmo-7B-Stage1, a Byte-Level Model Retrofitted from Olmo 3 7B

    AIAi2 has released Bolmo-7B-Stage1, a 7B byte-level language model retrofitted from Olmo 3 7B through a short additional training procedure. This checkpoint includes only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.