Skip to contentSkip to stories

Updated

#Model release

Aug 26

Aug 26Wed
  1. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  2. Ai2 · new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  3. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  4. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Z.ai Release NotesAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  3. Daniel HanAI score34

    Fine-tune Qwen3.8-27B free on Kaggle with Unsloth QLoRA

    AIDaniel Han says users can fine-tune Qwen3.8-27B for free on Kaggle with a Google account, which provides 30 hours of GPU time on 2× Tesla T4s. Using QLoRA and Unsloth's kernels, the 27B model fits within 24 GB VRAM with no accuracy loss, according to the post. The background post from Unsloth adds that its notebook trains Qwen3.8-27B 1.5x faster with 50% less VRAM.

  4. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  5. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging FaceAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    AIGoogle Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Qwen · new models on Hugging FaceAI score75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    AIQwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.

Aug 21

Aug 21Fri
  1. DeepSeekAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

  2. DeepSeek API NewsAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu

Aug 19

Aug 19Wed
  1. Liquid AI BlogAI score60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    AILiquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.