Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogOfficialAI score52

    Harvey Tenet, a legal model post-trained from Kimi K3 with Fireworks

    AIHarvey and Fireworks post-trained Tenet from the Kimi K3 base using asynchronous reinforcement learning on the Fireworks Training API for long-horizon legal work. On the Legal Agent Benchmark, Tenet reached 19.7% all-pass versus 10.8% for base Kimi K3, and its cost per task was $5.92 versus $5.62.

  3. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  4. Daniel HanXAI score34

    Fine-tune Qwen3.8-27B free on Kaggle with Unsloth QLoRA

    AIDaniel Han says users can fine-tune Qwen3.8-27B for free on Kaggle with a Google account, which provides 30 hours of GPU time on 2× Tesla T4s. Using QLoRA and Unsloth's kernels, the 27B model fits within 24 GB VRAM with no accuracy loss, according to the post. The background post from Unsloth adds that its notebook trains Qwen3.8-27B 1.5x faster with 50% less VRAM.

  5. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  6. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. Chip HuyenXAI score13

    Chip Huyen asks why GPT 5.6 over-engineers solutions

    AIChip Huyen asks why GPT 5.6 over-engineers so much, posing the question without offering an explanation or supporting data. The post is a brief user observation about the model's tendency toward excessive complexity.

  2. Google · new models on Hugging FaceOfficialAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    AIGoogle Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  3. ReplicateOfficialAI score12

    Alibaba's Wan 3 model now available to try on Replicate

    AIReplicate has made Alibaba's Wan 3 available to try through a hosted demo page. The post links to the model page at and provides no further details on capabilities, specifications, or pricing.

  4. Microsoft ResearchOfficialAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post
  5. GeneralistOfficialAI score27

    Generalist releases GEN-1.5, a foundation model for physical-world robotics

    AIGeneralist has announced GEN-1.5, its latest foundation model for the physical world. The post provides only a link to the company's blog for further details, so no specifications, benchmarks, or availability information can be confirmed from this source.

  6. GeneralistOfficialAI score38

    Generalist reduces time from physical prompt to robot behavior with GEN-1.5

    AIGeneralist says it has reduced the time needed to go from a physical prompt to robot behavior, making it faster to teach robots new tasks. The company links this speedup to easier scaling of physical work, and points readers to its GEN-1.5 blog post for details.

    Video from @GeneralistAI's post
  7. Meituan LongCatOfficialAI score23

    LongCat-2.0 now available in opencode Go for developers

    AIMeituan LongCat has made LongCat-2.0 available in opencode Go, according to the post. The post describes LongCat-2.0 as a 1.6T-parameter model with 48B active parameters, a 1M-token context window, and fully open-source release. Meituan LongCat invites users to try the model in opencode and share what they build.

  8. Qwen · new models on Hugging FaceOfficialAI score75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    AIQwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.

Aug 21

Aug 21Fri
  1. Sundar PichaiXAI score60

    Gemini 3.7 Flash Posts Fastest Early Growth for a Gemini Model

    AISundar Pichai says Gemini 3.7 Flash set new Gemini growth records in its first week, making it the fastest-growing Gemini model so far. The model is now running in Search and the Gemini app. A quoted ARC-AGI post reports 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12 per task.

  2. Thinking MachinesOfficialAI score27

    Thinking Machines offers Inkling and Inkling-Small on OpenRouter

    AIThinking Machines announced that its Inkling and Inkling-Small models are available to try on OpenRouter. The post links to the Thinking Machines provider page on OpenRouter, with no further details on specifications, benchmarks, or pricing.

  3. Z.aiOfficialAI score28

    ZCode and GLM-5.3 Offer 100M Free Tokens to 50,000 New Users

    AIZ.ai is extending its Build Week into an ongoing series, giving 50,000 new ZCode users 100M free GLM-5.3 tokens each. The offer runs until August 23 at 6 PM PT, with the linked event page noting that unused tokens expire when the event ends.

  4. Microsoft AIOfficialAI score11

    Microsoft AI invites users to try a new model on MAI Playground

    AIMicrosoft AI is inviting users to try an unnamed model on MAI Playground starting today. The post provides no model name, version, benchmark results, or availability details beyond a link to the playground.

  5. Microsoft AIOfficialAI score34

    Microsoft AI launches MAI-Image-2.6 for consistent image editing

    AIMicrosoft AI introduced MAI-Image-2.6, its latest image model, which can modify color, style, and visual details while maintaining consistency across iterations. The post positions it as a tool for bringing mood boards to life.

    Video from @MicrosoftAI's post
  6. DeepSeekOfficialAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

    Image from @deepseek_ai's post
  7. DeepSeek API NewsOfficialAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu

Aug 19

Aug 19Wed
  1. Liquid AI BlogOfficialAI score60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    AILiquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

  2. GeneralistOfficialAI score42

    Generalist's GEN-1.5 marks a new frontier in physical AI generality

    AIGeneralist says its GEN-1.5 model shows a new level of generality when pretrained on physical interaction data at a scale few thought feasible without shortcuts. The company says it does not yet see where performance gains level off.

  3. TinkerOfficialAI score38

    Qwen3.8-27B is now available on Tinker

    AITinker has made Qwen3.8-27B available today. The model is natively multimodal, handling images and video, with flexible thinking control. Tinker says it performs meaningfully better at coding, professional work, research, and long-horizon agentic tasks.

  4. Daniel HanXAI score40

    Unsloth releases 1-bit Qwen3.8-27B quants running on 8GB RAM

    AIUnsloth has released 1-bit quantized versions of Qwen3.8-27B that run on 8GB of RAM while retaining about 77% of BF16 accuracy. The team originally hesitated to publish them but was surprised by how well they performed in internal testing. The release accompanies new Qwen3.8-27B GGUFs that the company says deliver 10% higher accuracy.

  5. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  6. Google · new models on Hugging FaceOfficialAI score26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    AIGoogle has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

  7. Daniel HanXAI score40

    Unsloth releases Qwen3.8-27B GGUFs with Dynamic v3 quantization

    AIUnsloth released new Qwen3.8-27B GGUF quantizations built with Unsloth Dynamic v3, which it says gain about 10% top-1% accuracy at the same size. The accuracy was measured with the new Divergence-300 metric, which extends top-1% greedy accuracy to 32 tokens using 300 unseen examples from Terminal Bench and DeepSWE. Unsloth also released 1-bit quants that it says run in 6–8GB, with 8GB RAM cited for running them.

  8. Google · new models on Hugging FaceOfficialAI score22

    TIPS So400m/14 v1 Vision-Language Model Released on Hugging Face

    AIGoogle released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.

  9. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS L/14 v1 vision-language model on Hugging Face

    AIGoogle has published google/tipsv1-l14, the original v1 L/14 release of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The L/14 variant has 304M vision parameters and 184M text parameters at 448 resolution, with an embedding dimension of 1024, and is licensed under Apache 2.0.

  10. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS B/14 v1 vision-language model on Hugging Face

    AIGoogle has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.

Aug 18

Aug 18Tue
  1. Liquid AI BlogOfficialAI score65

    Liquid AI releases QAD 4-bit LFM2.5 checkpoints for edge deployment

    AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.

    Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.

  2. Microsoft AIOfficialAI score13

    Microsoft AI invites users to try a model on MAI Playground

    AIMicrosoft AI is inviting users to try a model today on its MAI Playground, linked in the post. The post, shared by the @arena account, does not name the model or give details about its capabilities.