Updated
#On-device
Updated
Oct 6
Philipp SchmidAI score22 Philipp SchmidAI score22 Try directly in your browser
vLLMAI score60 vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings
AIvLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.
Paige BaileyAI score54 EmbeddingGemma 2 launches as an Apache 2.0 multimodal embeddings model
AIGoogle's EmbeddingGemma 2 is an open embeddings model for on-device use that covers code, image, video, audio, and text. It comes in modular sizes from 270M text/code to 740M full multimodal, supports Matryoshka truncation down to 128 dimensions, and reports a 14% gain on MTEB Code over v1 under an Apache 2.0 license. The author's post highlights the release and a Hugging Face demo, while the benchmark table compares it with several models.
Unsloth AIAI score62 Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use
AIGoogle released EmbeddingGemma 2, a 740M-parameter open embedding model under Apache 2.0 that combines a 270M text model with vision (170M) and audio (300M) encoders. The 270M text model can run locally with 0.5GB of RAM, and the full multimodal model with 1GB, and Unsloth provides GGUF files and fine-tuning support.
GoogleAI score43 Best-in-class for its size at 740M parameters, EmbeddingGemma 2 outperforms some models over twice its size while using from as little as…
AI…191MB to 567MB of active RAM. It also features a 4x larger 8K context window than the first generation, processing up to 5.5 minutes of audio, 29 images, or 58 video frames in one pass.
GoogleAI score44 When paired with Gemma 4, EmbeddingGemma 2 powers efficient, on-device retrieval augmented generation (RAG) with a lower memory footprint.
AIEmbeddingGemma 2 retrieves local files, and Gemma 4 reasons over them for grounded answers with complete privacy.
Google for DevelopersAI score48 EmbeddingGemma 2 from @googlegemma, our compact 740M-parameter AI model built specifically for on-device and edge apps, has arrived.
AIWhile the first-generation model delivered best-in-class text embedding, EmbeddingGemma 2 natively understands it all, connecting text, images, audio, and video in one shared space. This means you can search across different formats like text, images, audio, and video, without having to manually organize, translate, or label them first.
Google for DevelopersAI score40 Google's multimodal embedding toolkit runs fully offline on device
AIGoogle's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.
Google DeepMindAI score58 Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0
AIGoogle DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.
Google DeepMindAI score62 Google DeepMind releases EmbeddingGemma 2, a natively multimodal open embedding model
AIGoogle DeepMind introduced EmbeddingGemma 2, its first natively multimodal open model for on-device embeddings. The model expands beyond text to unify code, images, audio, and video in a shared embedding space.
Sundar PichaiAI score62 Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use
AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.
Google DeepMind · The KeywordPickAI score72 Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use
AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.
Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.
IThome · AIAI score41 Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s
AIDeveloper Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.
Harrison ChaseAI score20 Good take on harnesses
Oct 5
Google Developers BlogPickAI score67 Google releases EmbeddingGemma 2, a multimodal embedding model for on-device search
AIGoogle DeepMind launched EmbeddingGemma 2, an open-weight 740M parameter model that maps text, images, video frames, and audio into one vector space. The model can run on-device, with about 567MB active RAM for the full multimodal model on a Google Pixel 11 Pro, and is available through Google AI Edge Gallery, Google AI Edge Foresight on Mac, and MediaPipe Tasks, with ML Kit support coming in the weeks ahead.
Why it matters: The post names concrete on-device apps, memory footprints, and latency figures, showing how a multimodal embedding model can power local search without cloud calls.
Liquid AI · new models on Hugging FaceAI score44 LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio
AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.
Oct 4
TekniumAI score29 We got ESP32 at home
Oct 3
Orange AIAI score55 Local Qwen Flash inference on consumer GPUs jumps roughly tenfold in a week
AIThe author reports that a dual RTX 5070 Ti setup running Qwen Flash rose from 200 prefill and 10 decode to 2200 prefill and 67 decode, now on a single card, using Strata and a custom PR. The post argues that such consumer-hardware speeds, once limited to top-end machines, could pressure the economics of selling model compute via API.
Oct 2
Aravind SrinivasAI score62 Perplexity open-sources models, an inference engine, and security tools
AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.
Georgi GerganovAI score34 Decision models in llama.cpp are now available The `/v1/systemone` endpoint is available in the latest llama builds.
AIUse it to do Jev-style inference locally, efficiently and privately. Multiple open models are supported with more to come.
NVIDIA BlogAI score43 NVIDIA DGX Spark 64GB Brings Local AI to More Developers at $4,999
AINVIDIA's DGX Spark 64GB configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Oct. 23, starting at $4,999. It supports models up to 100 billion parameters on device, and two units can be clustered via NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. NVIDIA says the clustered setup delivers up to 1.7x the performance of a single system in its Qwen 3.8 27B test.
Sep 30
OllamaAI score30 Ollama now supports Jev-like decision models all locally.
AIUse decision models like Nimble for tasks like ticket triaging, model routing, and content moderation. ollama pull nimble Here’s Nimble playing Ollama racer through the new local /v1/systemone API by making decisions in real-time. 🏎️
Sep 29
howie.seriousAI score23 Qwen3-VL 32B tags 40,000 Eagle images on a 128GB Mac
AIA user ran qwen3-vl:32b-instruct locally through Ollama on a 128GB computer to auto-tag 40,000 images in their Eagle app. The image library was collected over 10 years and is intended as a personal asset base for Claude Code-made knowledge videos. The user says the use case makes the 128GB memory purchase feel worthwhile.
Artificial Analysis ArticlesPickAI score62 Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents
AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.
Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.
Sep 28
Daniel HanAI score29 Unsloth Desktop can serve local Jev type decision models!
AIWe made a real time packing demo powered by local Laya through Unsloth’s Decision API - suitcase items update as you type and change your travel plans! More optims coming soon to make it even faster for local hardware!
Google Cloud · AI & Machine LearningAI score40 Why startups should pair open models like Gemma 4 with frontier APIs
AIGoogle Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.
OpenBMBAI score22 Community EXL3 4-bit quantization shrinks MiniCPM5-2B to 1.61 GB
AIA community-built 4.0 bpw EXL3 quantization of MiniCPM5-2B reduces the quantized model weights to 1.61 GB for local inference. The author reports roughly 68–70 tokens/s on an NVIDIA Tesla T4, and the model runs with ExLlamaV3 and TabbyAPI.
Sep 26
Liquid AIAI score20 What makes an on-device agentic model actually useful?
AIPost-training. @maximelabonne, @EdoardMosca, and Jiahui Wang from Liquid’s post-training team discuss how models learn to use tools, follow instructions, handle longer contexts, and recover when tasks get complex.
Sep 24
Google for DevelopersAI score37 Gemma 4 is now running locally on-device in the @Antigravity SDK.
AIBuild fully local or hybrid multi-agent workflows that pair cloud models with a @GoogleGemma 4 workforce. Audit, patch, and test your code with total data privacy and zero API fees—all powered by LiteRT under the hood.
Liquid AI NewsletterAI score38 Liquid AI's Liquid Context now optimized for Snapdragon NPUs; LFM Longevity models released
AILiquid AI announced its on-device Liquid Context layer is now optimized for Snapdragon processors using the Qualcomm Hexagon NPU, letting edge agents learn user routines and share context across devices. Separately, Liquid AI released LFM2-1.2B-Longevity and LFM2-2.6B-Longevity, which the company says often match or outperform much larger frontier LLMs on longevity prediction tasks.
OpenBMBAI score34 FIT-GGUF enables size-targeted mixed-precision quantization of MiniCPM5-2B
AIDeveloper @Scorp1o_117 used FIT-GGUF to build four MiniCPM5-2B GGUF variants, ranging from about 1.14 GiB to 1.46 GiB, tuned to target file sizes or fidelity tiers. Instead of fixed presets, FIT-GGUF allocates precision tensor by tensor, with Quality, Balanced, Compact, and Mini options, and its generated files matched predicted sizes. Builds are evaluated with KL Divergence and Same-top metrics and are available on Hugging Face.
Tencent HunyuanAI score34 🚀 Tencent Hy Translation just landed.Powered by Hy-MT2.
AI33 languages. 5 Chinese minority languages & dialects. Voice. Photo. Full offline — on-device, no network required. Already live in 12 countries and regions. Travel, drive, work or read abroad. Accurate. Natural. Always available. ⏬ ⏬⏬
Sep 23
Liquid AI BlogAI score46 LFM2.5-VL-DSpark speeds up vision-language model decoding on GPUs and edge devices
AILiquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, delivering decoding throughput gains of up to 2.66× on GPUs and 3.13× on edge devices. The drafter adds about 280M parameters, an 8.9% increase in the deployed model's parameter count, and is available on Hugging Face with support in llama.cpp, SGLang, and MLX-VLM.