Skip to contentSkip to stories

Updated

#On-device

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. The Verge · AINewsAI score40

    Alexa Plus excels at running a smart home but falls short as a personal assistant

    AIAmazon's Alexa Plus, powered by generative AI, now responds in three to five seconds and handles multistep smart home commands, cooking questions, and calendar imports more reliably than the original Alexa, according to a year-long test by The Verge. The reviewer says its personal assistant features remain underbaked and frustrating, and that ads on Echo Show displays are excessive. Alexa Plus costs $19.99 a month in the U.S. unless users have an Amazon Prime membership, and the Echo Dot Max is recommended as the ad-free option.

  2. OpenBMBOfficialAI score28

    MiniCPM5-2B runs at 37 tok/s on iPhone Air

    AIOpenBMB reports that its MiniCPM5-2B model runs at 37 tokens per second on an iPhone Air. The post presents this as evidence that small open multimodal models can run on mobile devices without a cloud GPU, with NobodyWho noting the model is available in its Chat app.

  3. ModelScopeOfficialAI score63

    Google releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device search

    AIGoogle released EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0 for private, on-device search and retrieval. It maps text, code, images, video, and audio into one shared space and reports a 9.92-point gain over EmbeddingGemma 1 on MTEB Code. The post lists about 191MB active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro.

    Video from @ModelScope2022's post

Oct 8

Oct 8Thu
  1. TechCrunch · AINewsAI score38

    Google launches Google AI Edge Foresight, a local-first Mac meeting note-taker rival to Granola

    AIGoogle released Google AI Edge Foresight, a Mac app that captures meeting notes offline using the on-device EmbeddingGemma 2 model with 740 million parameters. The app offers split-screen shorthand and AI-generated notes, transcripts, and a Gemma 4-powered assistant that can answer questions from uploaded documents. Google's FAQ says it is optimized for Apple Silicon.

  2. Microsoft CopilotOfficialAI score29

    Copilot on Windows gains Hybrid Intelligence, blending cloud and local agents

    AIMicrosoft Copilot announced Hybrid Intelligence for Windows, which balances cloud and locally run agents to handle tasks like file wrangling and workflows on the PC. According to Satya Nadella, Copilot can use context on the PC with user permission and draw on local models when appropriate. The feature is described as coming soon.

  3. SantiagoXAI score34

    Atomic Agent Desktop launches with Linux support on day one

    AIAtomic Agent Desktop is a local AI-first agent available for Mac, Windows, and Linux. It offers a 4x larger context window on local models using TurboQuant, and Atomic Fusion orchestrates cloud and local models to reduce costs.

  4. 🚨 AI News | TestingCatalogXAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

    Video from @testingcatalog's post
  5. Comfy BlogOfficialAI score34

    How I Generated Live Video with MiniMax H3 on a Single GPU

    AIA ComfyUI developer generated 15-second 448×256 video in 15 seconds or less on one RTX 5090 using MiniMax H3 with FastVideo's FastH3 V2 checkpoint in four sampling steps. The setup combined sparse attention, a smaller ClipProj text encoder, a pruned INT8 checkpoint, and a fused FP4 MLP, cutting VRAM needs from 80GB to under 30GB. The custom ComfyUI node is open source.

  6. Zhihao JiaXAI score62

    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post
  7. SiliconANGLE · AINewsAI score30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    AILiquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  8. The Verge · AINewsAI score52

    Google's experimental AI Edge Foresight transcribes meetings fully offline on Mac

    AIGoogle has released AI Edge Foresight, a free experimental note-taking app that transcribes meetings and audio files entirely offline on macOS. It runs on the on-device EmbeddingGemma 2 model and turns shorthand notes into polished notes based on the transcript. Google says files, meeting audio, and notes never leave the computer, and the app is currently optimized only for Macs with Apple Silicon.

  9. Nous ResearchOfficialAI score31

    Hermes Agent runs on ASUS ProArt RTX Spark PCs with local models

    AINous Research says Hermes Agent is now available on the new ASUS ProArt RTX Spark PCs, with MuseTree and ComfyUI integrations and local model support on up to 128 GB of unified memory. The post presents it as a creative stack that runs entirely on the user's own machine.

  10. IThome · AINewsAI score40

    Microsoft Confirms Copilot+ PC Brand Lives On, Runs 2 Trillion Local AI Inferences Monthly

    AIMicrosoft Windows and devices head Pavan Davuluri confirmed the Copilot+ PC brand has not been discontinued, saying more than 40% of commercial laptops are Copilot+ PCs shipping in tens of millions annually. He said these devices run over 2 trillion local inferences per month across search, image processing, and video calls. Microsoft plans to strengthen them through hybrid intelligence with local context, local actions, and local models.

Oct 7

Oct 7Wed
  1. Google Developers BlogOfficialAI score62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    AIGoogle's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    Why it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  2. Amazon ScienceOfficialAI score35

    Amazon's AI smart glasses guide delivery drivers hands-free to doorsteps

    AIAmazon has developed AI-powered smart glasses that guide delivery drivers from their van to the customer's doorstep without using their hands. Amazon applied scientist Yelin Kim will discuss the computer vision and edge AI behind the system at COLM 2026.

    Video from @AmazonScience's post
  3. 🚨 AI News | TestingCatalogXAI score34

    Microsoft brings hybrid local-cloud intelligence to Copilot for Windows

    AIMicrosoft is adding hybrid intelligence to Copilot for Windows, letting it use local PC context and local models for tasks. Per Satya Nadella's quoted post, Windows will route each task to local or cloud models, and Copilot will act on the user's behalf only with permission.

    Video from @testingcatalog's post
  4. Satya NadellaXAI score38

    Microsoft brings Hybrid Intelligence to Copilot on Windows

    AIMicrosoft is upgrading Copilot on Windows with Hybrid Intelligence, which lets it use context from the user's PC, take actions on the user's behalf, and run local models when appropriate. With the user's permission, the feature aims to add capability while stretching token usage further.

    Video from @satyanadella's post
  5. NVIDIA BlogOfficialAI score67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    AINVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

  6. Satya NadellaXAI score72

    Windows adds on-device agents, local coding models, and Hybrid Intelligence

    AIMicrosoft says Windows will bring unmetered intelligence to PCs, letting agents work securely on-device. The post lists MAI-Code-1.1 Flash, a 137B parameter coding model with a 256K context window optimized to run on PCs, and GitHub Copilot handoffs to local models. It also describes Hybrid Intelligence, which lets Copilot act on the PC and keep sensitive work local, and Code in Copilot for building software without cloud token spend, on devices such as Surface Laptop Ultra powered by NVIDIA RTX Spark.

    Image from @satyanadella's post
  7. Liquid AIOfficialAI score38

    Liquid AI's Open d1 models run on NVIDIA hardware with llama.cpp support

    AILiquid AI's Open d1 models run across NVIDIA DGX, RTX, and Jetson hardware, with day-one llama.cpp support for deployment anywhere. Measured one request at a time, the d1-3B model's single-question latency is 8 ms on an NVIDIA RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB, and 50 ms on Jetson Orin Nano.

  8. Liquid AIOfficialAI score30

    Liquid AI shows d1-3B running 10 live-camera demos, one forward pass per frame

    AILiquid AI built 10 live-camera demos for its d1-3B model, ranging from gesture-controlled games to content moderation, each using one forward pass per frame. In collaboration with NVIDIA Robotics, the company also showed d1-3B navigating an environment in Isaac Sim, served on a Jetson in a hardware-in-the-loop setup.

    Video from @liquidai's post
  9. Liquid AIOfficialAI score36

    Liquid AI releases d1-omni-600M, a 600M multimodal model for on-device tasks.

    AILiquid AI has released d1-omni-600M, an experimental 600M-parameter model that handles text plus image or audio input. It combines LFM2.5-Encoder-350M with vision and audio encoders and leads the company's text benchmark comparison on toxicity detection and paraphrase identification. The post suggests uses such as voice-command routing, on-device moderation, and intent classification.

    Image from @liquidai's post
  10. Liquid AIOfficialAI score52

    Liquid AI releases open-weight d1-3B and d1-omni-600M multimodal models

    AILiquid AI released Open d1, two open-weight multimodal models in its d1 decision model family. The d1-3B model supports text and vision, while d1-omni-600M supports text plus image or text plus audio. The source says the models are meant for real-time decision making across data centers, RTX workstations, and Jetson edge devices.

    Image from @liquidai's post
  11. Hugging Face BlogOfficialAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  12. Aravind SrinivasXAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

  13. 🚨 AI News | TestingCatalogXAI score41

    Google releases Foresight macOS app using Gemma 4 for voice notes

    AIGoogle released the Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma 2. The app can connect to Google Drive to build a knowledge graph and, when transcription is active, uses local Gemma 4 E4B or Gemma 4 12B models to transcribe voice notes into new documents. EmbeddingGemma 2 is an open-weight, Apache 2 licensed 740M-parameter multimodal embedding model with an 8K context window.

    Video from @testingcatalog's post
  14. Teknium 🪽XAI score36

    Community brings Hermes Gadget SDK to LilyGO, AIPI Lite, and old Android phones

    AIDevelopers are running Hermes on devices such as LilyGO watches, AIPI Lite, desk gadgets, and old Android phones after the Hermes Gadget open SDK and demo were released three days ago. The post credits @NousResearch and says more boards are landing on main through contributor PRs, with the SDK available on GitHub.