Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 26

Aug 26Wed
  1. Replit BlogOfficialAI score38

    Replit Launches Intelligent Model Routing, Auto-Matching Each Task to Suitable AI Model

    AIReplit has made Intelligent Model Routing available to all users, automatically matching each task with a model based on quality, speed, and cost. In the company's testing, the feature delivered the same output quality at 65% lower cost than the previous version of Max Mode. Enterprise administrators can restrict routing to a company-approved set of models.

  2. LM StudioOfficialAI score57

    GLM-5.3-Flash by Z.ai is now live in LM Studio

    AILM Studio announced that Z.ai's GLM-5.3-Flash, previously previewed as Ox Alpha, is available in LM Studio Bionic. The source says the model outperforms GLM-5.2 at 9-10x lower cost, supports image input, and is served from US-based servers with ZDR enabled by default.

  3. Michael TruellXAI score60

    Grok Bot opens to all Grok and Cursor subscribers

    AIGrok Bot is now available to all standard Grok and Cursor subscribers, with SuperGrok and Cursor Pro subscribers included. Cursor's Michael Truell says users are delegating tasks ranging from running small e-commerce businesses to testing production software. Weekly usage limits are also being reset for all users.

  4. Sundar PichaiXAI score42

    Google launches Gemini 3.5 Transcribe with 85+ language support

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that auto-detects over 85 languages and handles multiple speakers. It also supports custom vocabulary adaptation for specialized jargon. The API is available now in Google AI Studio and Gemini Enterprise.

    Video from @sundarpichai's post
  5. SpaceXAIOfficialAI score42

    Grok Voice models now power LiveKit voice agents with ZDR support

    AILiveKit announces that developers can build voice agents using Grok Voice models, with full ZDR support. LiveKit's example patient intake agent cascades Grok STT, Grok 4.3, and Grok TTS through LiveKit Inference in a single AgentSession, with no separate API key or billing.

  6. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  7. Unsloth AIOfficialAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    AIUnsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    Why it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

    Image from @UnslothAI's post
  8. LMSYS OrgOfficialAI score60

    SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

    AISGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

  9. GeneralistOfficialAI score22

    GEN-1.5 shows physical prompt steerability in a shared environment

    AIGeneralist AI demonstrated that GEN-1.5 behaves differently when given different physical prompts in the same environment. The post shares a simple example responding to a request from @JagdeepBhatia8 about whether the model follows physical prompts or just the most likely action for the scene.

    Video from @GeneralistAI's post
  10. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  11. Ai2 · new models on Hugging FaceOfficialAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  12. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  13. Ai2 · new models on Hugging FaceOfficialAI score39

    Ai2 releases Bolmo-1B-Stage1, a byte-level version of OLMo 2 1B

    AIAi2 has released Bolmo-1B-Stage1, a 1B-parameter byte-level language model retrofitted from OLMo 2 1B to process bytes rather than tokens. This checkpoint includes Stage 1 training only, with inner model parameters unchanged, and is available on Hugging Face under an Apache 2.0 license for research and educational use.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  3. catXAI score38

    Claude unifies memory across Chat and Cowork surfaces

    AIAnthropic's Claude now shares one memory across chat and Claude Cowork, so information saved once carries over between surfaces. Cowork tasks start from what Claude already knows from your chats, such as project context, preferences, or past clients, and users decide what the memory holds.

  4. Andrew NgXAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    AIOpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  5. v0OfficialAI score40

    v0 adds Vercel Connect for secure access to Slack, GitHub, and more

    AIv0 now lets users connect apps directly to Slack, GitHub, Notion, Salesforce, and other services through Vercel Connect. Vercel describes Connect as generally available, offering short-lived scoped access tokens, token and trigger observability, and RBAC with audit trails for 100+ services.

  6. Microsoft AI BlogOfficialAI score24

    Microsoft Marketplace adds intelligent discovery to help businesses find AI apps and agents

    AIMicrosoft Marketplace has made intelligent discovery generally available worldwide, letting customers describe business challenges in natural language and receive contextual recommendations and conversational comparisons. Early preview results show customers were 68% more likely to find solutions that met their needs and take the next step toward purchase.

  7. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  8. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon