Skip to content
TodayOct 8Thu13 items
  1. TechCrunch · AI38

    Google launches Google AI Edge Foresight, a local-first Mac meeting note-taker rival to Granola

    Google released Google AI Edge Foresight, a Mac app that captures meeting notes offline using the on-device EmbeddingGemma 2 model with 740 million parameters. The app offers split-screen shorthand and AI-generated notes, transcripts, and a Gemma 4-powered assistant that can answer questions from uploaded documents. Google's FAQ says it is optimized for Apple Silicon.

  2. Microsoft Copilot29

    Hybrid Intelligence is moving closer to the work already happening on your PC. From file wrangling to carrying out workflows, Copilot on Windows balances cloud and locally run agents to help you stay in the flow of work, securely. Coming soon. https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/

    Hybrid Intelligence is moving closer to the work already happening on your PC. From file wrangling to carrying out workflows, Copilot on Windows balances cloud and locally run agents to help you stay in the flow of work, securely. Coming soon. https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/

  3. Testing Catalog62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    Atomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

  4. Comfy Blog34

    How I Generated Live Video with MiniMax H3 on a Single GPU

    A ComfyUI developer generated 15-second 448×256 video in 15 seconds or less on one RTX 5090 using MiniMax H3 with FastVideo's FastH3 V2 checkpoint in four sampling steps. The setup combined sparse attention, a smaller ClipProj text encoder, a pruned INT8 checkpoint, and a fused FP4 MLP, cutting VRAM needs from 80GB to under 30GB. The custom ComfyUI node is open source.

  5. SiliconANGLE · AI30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    Liquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  6. The Verge · AI52

    Google's experimental AI Edge Foresight transcribes meetings fully offline on Mac

    Google has released AI Edge Foresight, a free experimental note-taking app that transcribes meetings and audio files entirely offline on macOS. It runs on the on-device EmbeddingGemma 2 model and turns shorthand notes into polished notes based on the transcript. Google says files, meeting audio, and notes never leave the computer, and the app is currently optimized only for Macs with Apple Silicon.

  7. Nous Research31

    Advancing our mission to deliver ubiquitous powerful agents to everyone, Hermes Agent is on the new ASUS ProArt RTX Spark PCs: MuseTree and ComfyUI integrations plus local model support on up to 128 GB of unified memory. A complete creative stack that runs on your own machine.

    Advancing our mission to deliver ubiquitous powerful agents to everyone, Hermes Agent is on the new ASUS ProArt RTX Spark PCs: MuseTree and ComfyUI integrations plus local model support on up to 128 GB of unified memory. A complete creative stack that runs on your own machine.

  8. Merve Noyan37

    new Llama.cpp release ships with (multimodal!) Jev-like models support, performance upgrade for Metal and more! 🔥 super simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here https://huggingface.co/models?apps=llama.cpp&other=decision-model&sort=trending we also polished Llama App website & docs https://llama.app 🌟

    new Llama.cpp release ships with (multimodal!) Jev-like models support, performance upgrade for Metal and more! 🔥 super simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here https://huggingface.co/models?apps=llama.cpp&other=decision-model&sort=trending we also polished Llama App website & docs https://llama.app 🌟

  9. IT之家 · 人工智能40

    Microsoft Confirms Copilot+ PC Brand Lives On, Runs 2 Trillion Local AI Inferences Monthly

    Microsoft Windows and devices head Pavan Davuluri confirmed the Copilot+ PC brand has not been discontinued, saying more than 40% of commercial laptops are Copilot+ PCs shipping in tens of millions annually. He said these devices run over 2 trillion local inferences per month across search, image processing, and video calls. Microsoft plans to strengthen them through hybrid intelligence with local context, local actions, and local models.

Oct 7Wed
  1. Google Developers Blog62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    Google's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    Why it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  2. Amazon Science35

    AI-powered smart glasses that guide delivery drivers from van to the customer's doorstep, hands-free. Hear from Amazon applied scientist @_yelinkim on the computer vision and edge AI behind it at @COLM_conf. #COLM2026

    AI-powered smart glasses that guide delivery drivers from van to the customer's doorstep, hands-free. Hear from Amazon applied scientist @_yelinkim on the computer vision and edge AI behind it at @COLM_conf. #COLM2026

  3. Satya Nadella38

    We’re supercharging Copilot on Windows with Hybrid Intelligence. With your permission, Copilot can tap into the context on your PC, take action for you, and use local models when it makes sense, giving you more capability while helping your tokens go further.

    We’re supercharging Copilot on Windows with Hybrid Intelligence. With your permission, Copilot can tap into the context on your PC, take action for you, and use local models when it makes sense, giving you more capability while helping your tokens go further.

  4. Georgi Gerganov31

    It is very cool to see llama.cpp on the big stage in todays Windows event! The software and hardware stacks are finally coming together. Our community has put a lot of hard work in the past years and it shows. Looking forward to more users embracing local AI.

    It is very cool to see llama.cpp on the big stage in todays Windows event! The software and hardware stacks are finally coming together. Our community has put a lot of hard work in the past years and it shows. Looking forward to more users embracing local AI.

  5. NVIDIA Blog67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    NVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

  6. Satya Nadella72

    Windows adds on-device agents, local coding models, and Hybrid Intelligence

    Microsoft says Windows will bring unmetered intelligence to PCs, letting agents work securely on-device. The post lists MAI-Code-1.1 Flash, a 137B parameter coding model with a 256K context window optimized to run on PCs, and GitHub Copilot handoffs to local models. It also describes Hybrid Intelligence, which lets Copilot act on the PC and keep sensitive work local, and Code in Copilot for building software without cloud token spend, on devices such as Surface Laptop Ultra powered by NVIDIA RTX Spark.

  7. GitHub31

    📣 Coming soon to GitHub Copilot: intelligent local model routing. We are excited to announce the next step in our Project HydraFusion vision. GitHub Copilot will soon be able to automatically route tasks to a local model when it's most suitable, helping you save on AI credits.

    📣 Coming soon to GitHub Copilot: intelligent local model routing. We are excited to announce the next step in our Project HydraFusion vision. GitHub Copilot will soon be able to automatically route tasks to a local model when it's most suitable, helping you save on AI credits.

  8. Liquid AI40

    Choose d1-3B for high-quality text and vision decisions, or explore our experimental d1-omni-600M when footprint matters. Download the weights, fine-tune, and deploy locally. > Blog: https://www.liquid.ai/blog/d1-open > d1-3B: https://huggingface.co/LiquidAI/d1-3b > d1-omni-600M: https://huggingface.co/LiquidAI/d1-omni-600M > Docs: https://docs.liquid.ai/lfm/models/decision-models

    Choose d1-3B for high-quality text and vision decisions, or explore our experimental d1-omni-600M when footprint matters. Download the weights, fine-tune, and deploy locally. > Blog: https://www.liquid.ai/blog/d1-open > d1-3B: https://huggingface.co/LiquidAI/d1-3b > d1-omni-600M: https://huggingface.co/LiquidAI/d1-omni-600M > Docs: https://docs.liquid.ai/lfm/models/decision-models

  9. Liquid AI38

    Both Open d1 models run across NVIDIA DGX, RTX, and Jetson, with day-one llama.cpp support to run anywhere. d1-3B single-question latency, measured one request at a time: > NVIDIA RTX 4090: 8 ms > Jetson AGX Thor: 16 ms > Jetson AGX Orin 64 GB: 26 ms > Jetson Orin Nano: 50 ms 4/

    Both Open d1 models run across NVIDIA DGX, RTX, and Jetson, with day-one llama.cpp support to run anywhere. d1-3B single-question latency, measured one request at a time: > NVIDIA RTX 4090: 8 ms > Jetson AGX Thor: 16 ms > Jetson AGX Orin 64 GB: 26 ms > Jetson Orin Nano: 50 ms 4/

  10. Liquid AI30

    We built 10 live-camera demos for d1-3B, from gesture-controlled games to content moderation, with one forward pass per frame. In collaboration with @NVIDIARobotics, we also show d1-3B navigating an environment in Isaac Sim, with the model served on a Jetson in a hardware-in-the-loop setup. 5/

    We built 10 live-camera demos for d1-3B, from gesture-controlled games to content moderation, with one forward pass per frame. In collaboration with @NVIDIARobotics, we also show d1-3B navigating an environment in Isaac Sim, with the model served on a Jetson in a hardware-in-the-loop setup. 5/

  11. Liquid AI36

    d1-omni-600M is an experimental 600M-parameter model for text + image or text + audio. It combines LFM2.5-Encoder-350M with vision and audio encoders, and leads our text benchmark comparison in toxicity detection and paraphrase identification. Use it for voice-command routing, on-device moderation, and intent classification. 3/

    d1-omni-600M is an experimental 600M-parameter model for text + image or text + audio. It combines LFM2.5-Encoder-350M with vision and audio encoders, and leads our text benchmark comparison in toxicity detection and paraphrase identification. Use it for voice-command routing, on-device moderation, and intent classification. 3/

  12. Hugging Face Blog49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    Liquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  13. Aravind Srinivas62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    Perplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

  14. Testing Catalog41

    ICYMI 👀: Google released a new Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma2 for recording voice notes. > The app can connect to your Google Drive in order to build a knowledge graph. > When transcription is active, it uses local models, Gemma 4 E4B or Gemma 4 12B, to transcribe voice notes and convert them into a new document.

    ICYMI 👀: Google released a new Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma2 for recording voice notes. > The app can connect to your Google Drive in order to build a knowledge graph. > When transcription is active, it uses local models, Gemma 4 E4B or Gemma 4 12B, to transcribe voice notes and convert them into a new document.

  15. IT之家 · 人工智能22

    Deepal L06 2027 model officially announced as long-range magnetorheological AI sedan with two "lobsters"

    Deepal Automobile officially announced the 2027 Deepal L06, positioned as a long-range magnetorheological AI sedan equipped with two "lobsters," namely an AI Agent and an AI Box. Chairman Deng Chenghao said the two systems let owners state their intent in a single sentence, though no launch date has been disclosed.

Oct 6Tue
  1. meng shao62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.

  2. Liquid AI Blog62

    Liquid AI releases open d1-3B and d1-omni-600M decision models for edge devices

    Liquid AI released two open-weight d1 decision models, d1-3B and d1-omni-600M, on Hugging Face. d1-3B scores 48.57 on the Decision Index v0.2.1 public split and answers a single question in 8 ms on an NVIDIA GeForce RTX 4090 and 50 ms on a Jetson Orin Nano. d1-omni-600M is an experimental checkpoint that handles text with images or audio and scores 15.95 on the same index.

    Why it matters: The release pairs open-weight decision models with measured latency across Apple, NVIDIA, and Jetson hardware, showing how edge deployment changes what is practical.

  3. Google DeepMind67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    Why it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.