Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 3

Sep 3Thu
  1. Awni HannunAI score51

    Mirai releases speculative decoding in Uzu for Qwen3.6-27B on Apple M5 Max

    AIMirai is releasing speculative decoding in its Uzu inference engine, starting with Qwen3.6-27B. The quoted post reports 105 output tokens per second on an Apple M5 Max with 128 GB of unified memory, 2.9× faster than the fastest MLX speculative-decoding implementation Mirai benchmarked. The stack combines DFlash with Mirai's Weaver model, tree-based speculative decoding, Mirai quantization, and Metal kernels for Apple silicon.

  2. Engineering at MetaAI score34

    Meta's ZGateway Proxy Unifies ZippyDB Client Traffic to Cut Connection Sprawl

    AIMeta has introduced ZGateway, a stateless proxy tier that now carries about 40% of all ZippyDB traffic, projected to exceed 60%, and handles over 1 billion operations per second. The proxy collapses the many-to-many client-to-database connection mesh into two bounded hops, adding about 6% computational overhead in an average use case. It also enables admission control, load balancing, and cross-region resilience, which contain reconnection storms that previously caused host crashes.

  3. Baidu Inc.AI score44

    Baidu and IFAW launch AI Guardian to combat illegal wildlife trade

    AIBaidu and IFAW have launched AI Guardian, a platform powered by ERNIE models that builds on a collaboration since 2020 that has helped remove more than 18,000 listings linked to the illegal wildlife trade. The platform aims to make this wildlife protection technology more accessible, and users can sign in through ai4wcp.com to help protect wildlife.

Sep 2

Sep 2Wed
  1. Gemini NotebookAI score42

    Gemini Notebook replaces per-artifact daily limits with flexible shared usage limits

    AIGemini Notebook is moving from fixed daily limits per artifact to flexible usage limits, with free-tier users getting 3x more Audio Overviews, 10x more Reports, and 20x more Quizzes and Flashcards over 24 hours. The post adds that most users should not hit their limits in a regular session, and deferred artifact generation lets them keep creating throughout the day. Limit resets are shifting to every 5 hours, according to the companion post.

  2. Google AI DevelopersAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post

Sep 1

Sep 1Tue
  1. Meituan LongCatAI score46

    LongCat-2.0 Now Free to Try in Cline

    AIMeituan's LongCat-2.0, a 1.6T open-weights MoE model with a 1M context window, is now free to use in Cline. Cline's post says it scores similarly to Claude Opus 4.7 and Gemini 3.1 Pro. Users can select it under free models via /model after installing Cline with npm i -g cline.

  2. Cursor ChangelogAI score62

    Cursor adds self-hosted machines that keep tool execution inside your network

    AICursor now supports self-hosted machines, so tool execution stays on your own infrastructure while the agent makes tool calls locally. Team pools are named worker queues that scale with requests and can hibernate idle machines, restoring them within a reconnect window. Cloud agents can also run on sandboxes such as AWS Lambda, Cloudflare, Modal, and Vercel, and self-hosted workers now support computer use on Linux and Mac.

    Why it matters: The update explains how self-hosted workers keep tool execution inside your network while pools scale and hibernate, which matters for teams with strict data controls.

  3. Google AI StudioAI score75

    Google adds agentic video understanding to Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite

    AIGoogle AI Studio says agentic video understanding is now available across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the Gemini API. The company reports cost reductions of up to 66%, token consumption reductions of up to 88% and accuracy gains of up to 7% on standard video benchmarks. Developers enable it by setting processing to "agentic" in the API configuration, at standard token pricing.

    Why it matters: The source gives concrete cost and token figures and explains how the agentic loop replaces fixed-rate frame ingestion, helping developers weigh it against their current video pipelines.

  4. Google AI StudioAI score62

    Google AI Studio introduces agentic video understanding with Gemini

    AIin which the model decides what to watch, at what speed, and through which modality. It fetches only the moments and signals it needs instead of ingesting media at a fixed frame rate. The post says this cuts costs by up to 66% and token consumption by up to 88% while boosting accuracy, and it is available now via the Gemini API and in AI Studio.

    Video from @GoogleAIStudio's post
  5. World LabsAI score38

    World Labs' Atlas reconstructs spaces from few photos for robot simulation

    AIWorld Labs says its Atlas model reconstructs a space from just a few photos and generates photorealistic RGB and depth data that a robot's sensors would observe on any trajectory. The company says this lets robots be trained and tested in far more spaces, since previously scanning such spaces required expensive equipment and time-consuming capture.

    Video from @theworldlabs's post
  6. Gemini API ChangelogAI score62

    Gemini API adds agentic video understanding for three Gemini models

    AIGoogle released agentic video understanding for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite across the Interactions and GenerateContent APIs. The model dynamically navigates video timelines, requesting transcripts, frames, or audio tracks on demand. The source says this approach uses up to 88% fewer tokens for long-form content than static processing.

    Why it matters: The changelog names the affected models and API surfaces, and states a token-use figure that helps developers judge the cost of long video workloads.

Aug 31

Aug 31Mon
  1. The Register · AIAI score55

    OpenClaw 2.0 simplifies setup and adds shared sessions, but security defaults remain weak

    AIOpenClaw 2.0 is an open-source, self-hosted AI agent harness whose update simplifies installation, rebuilds the browser interface, and adds shared cloud sessions for multiple users. The article says the patch notes state shared session controls are not a security boundary, secret store values are not encrypted at rest, and sandboxing is off by default.

  2. Liquid AI NewsletterAI score46

    Liquid AI launches Pipette, an open-source benchmark for on-device foundation models

    AILiquid AI and Artificial Analysis released Pipette, an open-source benchmark platform for foundation models on edge devices, covering over 1,000 configurations across 30+ models. It measures five on-device metrics, including throughput, latency, context scaling, and memory use, on macOS, Windows, iOS, and Android. Liquid AI also said its updated LFM2.5 Q4_0 checkpoints, trained with Quantization-Aware Distillation, retain roughly 97% of BF16 baseline performance and suffer 73.4% less quality loss than standard post-training Q4_0 quantization.

Aug 30

Aug 30Sun
  1. Fireworks AI BlogAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  2. Chips and CheeseAI score62

    IBM explains its dual-ISA z/Architecture and Arm core design at Hot Chips 2026

    AIIBM's Christian Zoellin and Christian Jacobi discuss a next-generation processor that supports both z/Architecture and Arm instruction sets at Hot Chips 2026. Zoellin says separate decoders sit in a shared decode pipeline, while the caches, TLBs, and register files are reused, and endianness is handled in the load-store unit. Jacobi explains that Arm support aims to bring its software ecosystem to mainframe workloads, and that Spyre's memory bandwidth needs grew as use cases shifted toward agentic AI.