Granola now ships HIPAA compliance for Enterprise plans
AIGranola announces that HIPAA compliance is now shipping for its Enterprise plans. Customers interested in details are directed to contact the sales team.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIGranola announces that HIPAA compliance is now shipping for its Enterprise plans. Customers interested in details are directed to contact the sales team.
AIMeta has open-sourced Rebalancer, an assignment-problem solver it has used for over nine years to allocate resources across its infrastructure. The library separates problem specification, in-memory storage, solving, and debugging, and translates problems into expression graphs solved via local search or mixed integer programs using FICO Xpress, Gurobi, or the open-source HiGHS solver.
AIGoogle's Gemini Notebook says Live Chat is now fully rolled out to all Ultra subscribers on mobile. The feature enables real-time voice conversations with notebooks in about 100 languages, powered by the latest audio models. Users can ask questions about their sources and receive step-by-step guidance largely hands-free.
AIDeveloper Mark Fenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. The model selects key passages from official AI and technology sources, and a rule-based editorial layer preserves dates and context before assembling a daily recap. Invalid or incomplete outputs are rejected, and the full pipeline runs locally without a hosted-model fallback.

AIA research-agent design has GPT-6 orchestrating while MiniCPM5-2B runs locally on DGX Spark to read sources and extract traceable evidence. The post presents this as an example of pairing frontier reasoning with compact, efficient models for transparent research workflows.
AIA community developer integrated the Jev method into MiniMax H3 to sparsify attention, deciding per layer which parts to keep. On an RTX 4070, video generation time dropped from 6 min 7 sec to 3 min 34 sec, a 41.7% reduction. The post notes that Jev selected sparsity rates of 1%, 3%, 5%, and 10% across 49 layers in 4-step generation.
AIModelScope now supports both inference and LoRA training for Qwen-Image-2.1 on its Civision platform. The post directs users to to start creating.

AIWorkBuddy has added GLM-5.3-Flash to its model lineup and made it available now on the platform. The post invites users to try the model on their next task, but gives no further details on capabilities, speed, pricing, or benchmarks.

AIOpenMed reports that MiniCPM5-2B, run on a local Mac, correctly kept naproxen in medication history rather than the current-medication export after a newer note said it was stopped. Every graph connection in the run links back to its source, using fictional clinical notes.
AISGLang-Diffusion now serves Qwen-Image-2.1 for text-to-image generation, multi-image editing, and transparent RGBA output. The quoted SGLang post reports 1024×1024 generation in 18.7s and editing in 21.7s on a single RTX 4090 24GB with CPU offload, using 22.7 GiB peak GPU memory, with no quantization.
AIQwen-Image-2.1 is now available as a live Hugging Face Spaces demo, letting users try image generation and editing in the browser without setup. The demo runs on a single checkpoint that handles both tasks. Background notes describe a 7B-parameter model supporting up to 10 image references, with integration into diffusers and ComfyUI.
AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.
AIQwen-Image-2.1 from Alibaba's Qwen has day-0 support in vLLM-Omni, enabling one model to generate, edit, and create transparent images. The model pairs a 7.1B DiT with Qwen3-VL-8B. Setup details and supported combinations are available in the vLLM recipes page.
AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

AIQwen-Image-2.1 supports several ways to specify local edits, including circles marking regions on an image. In the example, the model removes a metal watch, changes hair color to black, and replaces clothing in three circled areas in a single pass.

AIQwen-Image-2.1 now generates transparent images natively, without post-processing. It also supports editing transparent images directly.

AIModelScope has added inference and training support for Qwen-Image-2.1 on its Civision platform. The feature is currently available on the CN site, with international support promised as coming soon.
AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.
AIWan 3.0 lets users place their own photo into any scene and animate it using the Peel-Off Sticker skill on wan.video. The post presents this as a single-photo workflow but gives no further technical details.
AIThe main post suggests ChatGPT Atlas, an AI browser, is being revived within the Codex and ChatGPT app. It also speculates that Codex or Claude Code built-in browsers could one day serve as full everyday browsers.
AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.
AILM Studio users can enable the Splash engine under Settings > Runtime > Experimental backends and download supported models by searching "incoai." The post recommends an M3 or newer Mac running macOS 26.4 with 36GB+ RAM.

AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.
AIMeta is opening access for developers to build connectors for Muse, its agent platform. Developers supply the API, while Muse provides the agent, browser, and user context, so people can reach a service simply by asking and their agent handles the rest. New connectors are live today at
AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.
AIChatGPT now supports connecting multiple accounts to most plugins, so users can bring work and personal context into one conversation. Connections are made in the plugin directory at Developers need no changes, though adding a profile tool to their MCP server lets ChatGPT label each account.
AIClaude Code version 2.1.277 now reads AGENTS.md when a folder has no CLAUDE.md, according to Anthropic engineer Thariq Shihipar's post. The behavior can be toggled in /config. Thomas Dohmke's main post jokingly says AI is finally aligned, with no further technical detail.
AIGoogle AI Developers demonstrates Gemini 3.8 Live Extended Thinking as a workshop assistant that analyzes a user's workspace, reasons aloud, and guides a beginner through building a cyberdeck. The demo targets users with no prior experience building one.
AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.
AIxAI's Grok Voice Transcribe 2.0 is twice as accurate as its predecessor in customer-support calls, spoken credentials, and short voice commands. The post lists these three scenarios as the areas of improvement, with no further figures given.

AIAtlassian reports that Transcribe 2.0 in Loom captured user instructions more accurately than its predecessor. Users can dictate change requests in Loom and export them directly to Cursor for coding.
AIxAI has made Transcribe 2.0 available today in the Grok Voice API. Pricing is $0.10 per hour for batch transcription and $0.20 per hour for streaming. The linked blog post contains further details.
AIxAI has introduced Grok Voice Transcribe 2.0, which it calls the world's most accurate speech transcription model. The post gives no benchmark scores, pricing, or availability details.

AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.
AIMiniMax has opened its Code CLI, making it available on GitHub at The post provides only the repository link and offers no further details on features or availability terms.
AILovable has acquired Sutro, a company that built a language, compiler, and backend platform making application behavior and rules explicit. Sutro founder Tomas Halgas and three engineers, Max Gfeller, Tony Zhan, and Hirad Arshadi, are joining Lovable, with Halgas leading its technical evangelism.
AIHuawei unveiled the Atlas 960E superpod, which links up to 4,096 NPUs using near-packaged optics (NPO) and claims eight exaflops at FP8 precision and up to one petabyte of high-bandwidth memory. Its Hi-ONE engine provides 7.2 terabits per second of transmission capacity, and Huawei says 5,500 engines replace 48,000 conventional 800G optical modules, cutting power consumption by more than 550 kilowatts. Huawei has proposed its NPO implementation agreement to the Optical Internetworking Forum, though it has not yet become a standard.
AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.