Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 11

Aug 11Tue
  1. Philipp SchmidBlogAI score60

    Gemini API now combines Google Search and Google Maps in one call

    AIGoogle Search and Google Maps can now be used in the same Gemini API call with Gemini 3.5 Flash and 3.6 Flash. Custom functions and MCP servers can be added to the same request through Tool Combination. The author says Gemini handles the search, place lookup, and function call in one interaction without extra roundtrips from the developer's side.

  2. IdeogramOfficialAI score32

    Ideogram launches Ad Resizer for multi-placement campaign exports

    AIIdeogram has introduced Ad Resizer, which takes one uploaded design and exports a campaign-ready set for placements including Instagram, Facebook, YouTube, connected TV, display, and print, even at extreme aspect ratios. The tool is available now in the Ideogram app and via the API.

    Image from @ideogram_ai's post
  3. Ali GhodsiXAI score42

    Databricks acquires ElectricSQL, the team behind PGlite

    AIDatabricks has acquired ElectricSQL, the team behind PGlite, a WebAssembly implementation of Postgres that runs in the browser and syncs asynchronously with Postgres instances. Databricks says this suits fast AI agents and plans to use the capabilities to strengthen its Lakebase Postgres offering.

  4. Kimi.aiOfficialAI score38

    Kimi K3 now available on Databricks via Unity AI Gateway

    AIMoonshot AI's open-weight model Kimi K3 is now live on Databricks through Unity AI Gateway. Customers can run it alongside their Lakehouse data with enterprise access controls, call monitoring, and performance tracking. The gateway lets teams test Kimi K3 against other frontier models without changing application code.

    Image from @Kimi_Moonshot's post
  5. Z.aiOfficialAI score33

    ZCode reaches 1 million users and resets GLM Coding Plan limits

    AIZ.ai says its ZCode platform has reached 1 million users, and it has reset usage limits for all GLM Coding Plan users as a thank-you. The post also announces an update aimed at turning long-horizon capabilities into completed engineering work, reporting a 98% cache hit rate that provides around 1.8x more usage.

    Video from @Zai_org's post

Aug 10

Aug 10Mon
  1. Jensen HuangXAI score62

    NVIDIA partners with six financial firms to mobilize $500 billion for AI infrastructure

    AINVIDIA says it has partnered with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to set up independent financing platforms aiming to mobilize over $500 billion of third-party capital for AI infrastructure. The article argues AI factories are investable assets because they generate revenue, serve many customers and improve through CUDA software. NVIDIA says it may provide residual-value support for up to 25% of an opportunity on a project-by-project basis.

  2. Liquid AI · new models on Hugging FaceOfficialAI score43

    LiquidAI LFM2.5-2.6B-DSpark Speeds Up LFM2.5 Decoding With Speculative Drafting

    AILiquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face. In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x. Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.

  3. Liquid AI · new models on Hugging FaceOfficialAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  4. Liquid AI · new models on Hugging FaceOfficialAI score42

    LiquidAI LFM2.5-1.2B-Instruct-DSpark Drafter Speeds Up Decoding About 2x

    AILiquid AI released LFM2.5-1.2B-Instruct-DSpark, a 295.7M-parameter speculative-decoding draft model for the LFM2.5-1.2B-Instruct target on Hugging Face. On an H100 it averages 4.81 accepted tokens per step and runs about 2.10x faster across benchmarks, with about 2x speedup in SGLang and on-device Apple silicon support via Metal.

  5. Andy JassyXAI score38

    Novo Nordisk selects AWS as preferred cloud and strategic AI partner

    AINovo Nordisk has chosen AWS as its preferred cloud provider and strategic AI partner to accelerate drug discovery. The collaboration will combine Novo Nordisk's scientific expertise with AWS AI tools, including Amazon Bio Discovery and Bedrock AgentCore, and establish a co-innovation hub in London. The partnership already spans AWS, Amazon Pharmacy, and One Medical.

    Image from @ajassy's post

Aug 9

Aug 9Sun
  1. Fireworks AI BlogOfficialAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

  2. PromptArmor Threat IntelligenceOfficialAI score65

    Malicious Zoom AI Skill Can Keep Attacker Connected and Exfiltrate Data

    AIPromptArmor reports that a malicious Skill or indirect prompt injection can make Zoom's ZoomMate agent connect to an attacker's server and run commands. The connection can persist after the user clicks stop or closes Zoom, and the final chat output appears normal.

    Why it matters: The report shows how a malicious skill or prompt injection can keep a Zoom agent connected after the user stops it, a risk to weigh before enabling agentic assistants.

Aug 8

Aug 8Sat

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceOfficialAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. Matei ZahariaXAI score44

    Matei Zaharia says AI Gateways let teams cut token costs centrally

    AIMatei Zaharia argues AI tokens are now a resource to optimize in software engineering, with companies routing all AI usage through an AI Gateway. The approach enables centralized analysis, which found settings on Claude Code and Codex that can substantially lower cost, plus smart routing and per-task budgets for engineers.

  3. Ali GhodsiXAI score58

    Databricks details four techniques it used to cut internal AI coding spend by up to 90%

    AIDatabricks published an analysis of four techniques it used to reduce internal AI spend while growing adoption, with savings of up to 90% in some scenarios. The techniques are shifting defaults to cheaper models such as GLM, automated task-level model routing, per-user spend visibility with adaptive budgeting, and pruning context bloat. The author, Ali Ghodsi, reposted Databricks co-founder Patrick Wendell's summary and recommended it.

  4. Ahmad Al-DahleXAI score23

    Airbnb says AI-native smaller teams launch concepts 60% faster

    AIAirbnb says its AI-native approach with smaller teams cut concept-to-launch time by up to 60% and nearly doubled shipped output, with about 80% more shipped in H1 than last year. The company says AI is now simply how it builds products.

  5. Prime Intellect BlogOfficialAI score62

    Prime Intellect adds multi-agent training and evaluation to PRIME-RL

    AIPrime Intellect's RL stack now supports multi-agent systems, letting users program interactions between agents, choose which roles learn, and assign credit across an episode. The release introduces Agent and Env abstractions and four example patterns: agentic judging, self-play, and user simulation. Multi-agent support ships today in verifiers 0.3.0 and prime-rl 0.8.0.

    Why it matters: The post explains the Agent and Env abstractions and four multi-agent patterns, showing how roles, credit assignment, and episodes can be programmed in one RL stack.

Aug 6

Aug 6Thu
  1. Lisa SuXAI score47

    AMD agrees to acquire AI inference startup Taalas

    AIAMD has agreed to acquire Taalas, an AI inference company, according to AMD CEO Lisa Su's post on X. Su described the Taalas team as phenomenal and working at the cutting edge of AI inference. The post gives no deal terms or closing timeline.

    Image from @LisaSu's post
  2. Intern Large ModelsOfficialAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

    Image from @intern_lm's post

Aug 5

Aug 5Wed
  1. Prime Intellect BlogOfficialAI score75

    Prime Agent launches open-source self-improving RLM coding harness

    AIPrime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.

    Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.

Aug 4

Aug 4Tue
  1. Fireworks AI BlogOfficialAI score26

    Voyage AI's embedding and reranking models now run natively on Fireworks AI

    AIVoyage AI by MongoDB's full lineup, including the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5, now runs natively on the Fireworks inference platform. The partnership lets teams run embedding, retrieval, reranking, and generation on one platform and one API. Fireworks says Voyage 4 Large outperforms Voyage 4, Voyage 4 Lite, Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large on average retrieval quality.

  2. Zed BlogOfficialAI score65

    Zed Enables OS-Level Sandboxing by Default for Its Agent Panel

    AIZed's agent panel now sandboxes its terminal and fetch tools by default, starting in release 1.14, and the restrictions are enforced by the operating system rather than by agent instructions. By default the sandbox blocks writes outside project directories, writes to .git, and network requests, and agents can request temporary escalation with a stated reason. The post also notes that sandboxing covers only those tools and does not protect against other tools, external programs, or the regular built-in terminal.

    Why it matters: The post explains how OS-enforced sandboxing limits agent terminal and fetch access, and why fine-grained command rules fall short of it.

  3. Mckay WrigleyXAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

Aug 3

Aug 3Mon
  1. JetBrains AI BlogOfficialAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

  2. Manus BlogOfficialAI score38

    Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps

    AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.

Aug 2

Aug 2Sun
  1. OpenRouter BlogOfficialAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Aug 1

Aug 1Sat
  1. Werner VogelsXAI score22

    Werner Vogels praises conversation with Clare Liguori on Kiro and agent support

    AIWerner Vogels called his conversation with Clare Liguori an excellent discussion of developer support for agents and Kiro. The quoted InfoQ podcast covers moving agents from demo to production, including why extra if statements can hurt agent performance, achieving high accuracy and low cost with small models, and observability within agent hops.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceOfficialAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. SkyworkOfficialAI score35

    Skywork AI Hardware Family's first Skywork Note batch sells out in one week

    AISkywork's first batch of its Skywork Note AI hardware device sold out one week after launch, prompting an accelerated rollout of the wider family, including the recording clip, the Recall pendant, and the TriRing AI ring. The company says the device is meant to capture real-world conversations and moments outside the screen, so users spend less time typing and more time away from it.