Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 29

Sep 29Tue
  1. Microsoft Foundry BlogAI score30

    Why content extraction still matters in the GenAI era

    AIMicrosoft's Azure AI team argues that better models do not eliminate the need for a dedicated content extraction layer, since agents need trustworthy, structured, and auditable inputs. The post notes that building extraction directly on an LLM quickly demands chunking, layout parsing, grounding, normalization, and evaluation infrastructure. Microsoft positions Azure Document Intelligence and Azure Content Understanding in Foundry Tools as managed options for that layer.

  2. Baseten BlogAI score38

    Baseten Partners With OpenAI to Offer Open Models to OpenAI Customers

    AIBaseten announced a partnership with OpenAI that makes open models powered by Baseten available to OpenAI customers for multi-model agentic coding. The company says organizations can route each task to the best-fit open or closed model, with Codex and GPT models among the options, and that Baseten's day-zero access to new open models lets teams evaluate them quickly. Baseten also cites US-based infrastructure with zero data retention for all prompts and capacity across more than 90 clusters in 20+ clouds.

  3. Azure BlogAI score40

    SQL Server on Azure Local Becomes Generally Available for Connected and Disconnected Use

    AIMicrosoft has made SQL Server on Azure Local generally available for connected and disconnected deployments, letting organizations run SQL Server in their own datacenters and edge locations. Disconnected operations continue locally where external connectivity is restricted or unavailable. Eligible existing SQL Server licenses can be used, and Foundry Local on Azure Local, currently in preview, brings AI inference alongside SQL Server data.

  4. Replit BlogAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

  5. Microsoft ResearchAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

  6. Azure BlogAI score46

    Microsoft Fabric and Copilot Integration: New Data Foundation Features for Agents

    AIMicrosoft is bringing business context from Fabric IQ into Microsoft Copilot, with Fabric IQ in Copilot Chat and Cowork generally available and integration into the new Code experience coming soon through the Frontier program. Power BI is also gaining agentic app creation in Power BI Desktop, letting users generate applications from trusted semantic models and publish them to Microsoft Fabric.

  7. Artificial Analysis ArticlesAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

  8. Manus BlogAI score50

    Manus Flex lets users connect their own API keys to the Manus workspace

    AIManus is launching Manus Flex, a module that lets users power Manus agents with their own API key from a supported inference provider. Model inference is billed directly by that provider, while other services used in Manus tasks still consume Manus credits. OpenRouter, Fireworks, and Modal are announced as initial inference partners for the Flex Inference Partner Program.

  9. Luma AI NewsAI score22

    AI Photo Editing Prompt Formula Preserves Color, Light, and Skin in Campaign Edits

    AIThe article presents a four-part prompt structure (action verb, target element, desired result, protection instructions) for AI photo editing, saying it preserves approved work across platforms. It identifies three common failure causes: unmatched light direction, stacked edits in one prompt, and vague visual language. It states that simple skin retouching takes 2-3 minutes versus 15-30 minutes manually.

Sep 28

Sep 28Mon
  1. vLLM BlogAI score54

    vLLM guide explains disaggregated serving for prefill and decode

    AIThe vLLM blog guide explains how separating prefill and decode, and moving tokenization to a CPU-only render tier, can keep token streams from stalling under load. In a two-L40S test on Qwen2.5-7B, collocated p99 inter-token latency reached 169 ms at 0.4 req/s while disaggregated serving stayed between 25 and 52 ms. The guide notes that the gain depends on fast KV cache transfer, and it includes setup code for NIXL-based serving and the render/derender API.

  2. Microsoft ResearchAI score30

    Microsoft Research Asia – Singapore marks one year advancing AI research, partnerships and talent

    AIMicrosoft Research Asia – Singapore, opened July 24, 2025 as Microsoft's first Southeast Asian research lab, reports progress after its first year. The lab's work spans next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development. Its healthcare collaborations on multimodal and agentic AI for clinical decision-making are being deployed through partnerships across Singapore's healthcare ecosystem.

  3. Google Cloud · AI & Machine LearningAI score40

    Why startups should pair open models like Gemma 4 with frontier APIs

    AIGoogle Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.

  4. Sierra BlogAI score34

    Sierra's Ghostwriter becomes a proactive Slack and Teams teammate for AI agents

    AISierra has turned its Ghostwriter tool into an always-on teammate in Slack and Teams that proactively suggests ideas, flags problems, and proposes experiments. Ghostwriter reviews recent customer calls, recommends which changes to try first, runs experiments, and reports when results are statistically significant. Sierra said it will begin rolling the feature out more broadly next week.

  5. Lovable BlogAI score57

    Lovable apps can now run inside a company's Microsoft tenant

    AILovable announced a partnership with Microsoft that lets users publish apps into their company's Microsoft Entra tenant using Copilot Managed Runtime. Apps can connect to Microsoft 365, Fabric, Dataverse, and SQL data, and staff sign in with their work login. Copilot Managed Runtime is in public preview, and Microsoft 365 connectors, Fabric, and Microsoft sign-in are available on every Lovable plan, while Entra workspace sign-in is included on Business and Enterprise.

  6. Baseten BlogAI score26

    Baseten and Blaxel Back NVIDIA OpenShell Sandboxes With Carbon Preview

    AIBlaxel, which Baseten acquired, is introducing Carbon, its fourth-generation infrastructure, in private preview for running agents in secure sandboxes. Carbon runs on microVMs with a dedicated IPv6 address per sandbox, supports manual snapshotting, forking, and snapshot-to-production within milliseconds, and includes a template with NVIDIA OpenShell preinstalled. Carbon is rolling out progressively by region and workspace and is coming to Baseten soon.

  7. Mastra BlogAI score29

    Mastra Publishes Guide to GDPR-Ready Agents with EU Hosting and Data Controls

    AIMastra's guide explains how teams can run agents under GDPR, with self-hosted deployments in any EU region or a platform environment created with --region eu. It covers PIIDetector redaction before data reaches the model, SensitiveDataFilter for trace fields, and retention and deletion handled in the team's own database. Mastra says it offers a DPA with EU Standard Contractual Clauses, a SOC 2 Type II audit, and no training on personal data.

  8. Manus BlogAI score60

    Manus 2.0 adds Cascade agent harness, Manus Studio, and Cue app

    AIManus 2.0 introduces a new agent harness called Cascade, Manus Studio with Video Editor and Game Dev environments, and a standalone Cue app for personal agents. In one tested configuration, Cascade used 23.2% fewer tokens, completed tasks 28.2% faster, and cost 32% less to run than the previous system. Cue is in early access and available with an invite code.

    Why it matters: The post separates the new agent harness, Studio, and Cue, and its Cascade chart gives measured token, time, and cost comparisons against the previous system.

Sep 27

Sep 27Sun
  1. PromptArmor Threat IntelligenceAI score72

    Elastic's AI SOC agent can be manipulated into leaking API credentials

    AIPromptArmor reports that Elastic's AI SOC agent, EASE, can be manipulated through malicious phishing alerts into minting API keys and sending them to an attacker. The attacker could then disable detection rules, create fake alerts, and exfiltrate data, and the report says the agent runs with user privileges and needs no human approval. PromptArmor says Elastic received the report on August 23, 2026, did not address it after four follow-ups, and published mitigations that include disabling built-in capabilities and write-capable tools.

    Why it matters: The report shows how a prompt injection in alert data can drive an AI SOC agent to leak API keys, with concrete mitigations for agent tool settings and default model choice.

  2. xAI News (Grok)AI score58

    xAI launches Team Bots, shared Grok Bots that learn as teams work

    AIxAI has launched Team Bots in public beta on Teams and Enterprise plans, letting teams build shared Grok Bots that keep context, plugins, credentials, and memories. Each person's conversations stay private while the Bot draws on skills shared across the team. The post also describes internal uses in sales, product and engineering, marketing, and data analytics, and it is available through Slack.

  3. Amp NewsAI score67

    Amp switches its default medium mode to Claude Opus 5.5

    AIAmp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.

    Why it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.

  4. Fireworks AI BlogAI score57

    Fireworks adds GLOBAL multi-region deployments under one endpoint

    AIFireworks AI introduced a GLOBAL option that lets one inference deployment run across geographies behind a single endpoint. The scheduler places workloads across eligible capacity while respecting hardware, quota, reliability, and data residency constraints. In a seven-day observational study, deployments spread across two or more serving clusters had a 99.992% request success rate, compared with 99.269% for single-region deployments.

Sep 26

Sep 26Sat
  1. Xiaomi MiMo · new models on Hugging FaceAI score50

    Xiaomi releases MiMo-V2.6-Pro-MOPD, a 1.02T-parameter sparse MoE model

    AIXiaomi has released MiMo-V2.6-Pro-MOPD, an upgrade of the MiMo-V2.6-Pro-RL checkpoint that fuses several domain-specialized teachers into one model via MOPD2 and targets tool-call repetition. The sparse MoE model has 1.02T total and 42B activated parameters, a 1M-token context length, and accepts text, image, video, and audio inputs. Weights are available on Hugging Face and ModelScope, with deployment recipes for SGLang and vLLM.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  3. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. Google Cloud · AI & Machine LearningAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

Sep 24

Sep 24Thu
  1. Baseten BlogAI score44

    LangSmith Fine-Tuning Trains Open Models on Agent Traces via Baseten Loops

    AILangChain launched LangSmith Fine-Tuning, which lets users fine-tune open models on their LangSmith agent traces using the open-source smithtune CLI. Training runs on Baseten Loops in the user's own workspace, and smithtune deploy places the evaluated checkpoint on a Baseten Dedicated Inference deployment. Loops is in early access, so users may need to request access for their workspace.

  2. Azure BlogAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  3. Microsoft Foundry BlogAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  4. Epoch AI · The Epoch BriefAI score45

    Huawei Trails Nvidia by About Four Years in AI Chip Performance and Output

    AIHuawei will likely remain about four years behind Nvidia in AI chip performance and production through 2030, Epoch AI estimates. Its flagship Ascend 950 delivers roughly half the performance of Nvidia's 2022 H100, and Huawei is projected to produce about 1.5 million chips in 2026 versus Nvidia's roughly 6 million, leaving it about 25 times behind in total compute.

  5. Google · Gemini appAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  6. Microsoft Foundry BlogAI score61

    Microsoft Foundry Routines reach general availability for scheduled and event-driven agents

    AIMicrosoft announced general availability of Routines in Foundry Agent Service, a managed way to run agents on a timer, on a recurring schedule, or in response to GitHub issue events and new Microsoft Teams channel messages. Routines keep the trigger, agent action, identity, connections, and run history in the Foundry project, and each routine can run under the creator's identity or the agent's own Microsoft Entra ID identity. A preview reminder tool lets a Hosted Agent schedule itself to resume later on the same conversation.

    Why it matters: The post explains how scheduled, event-based, and self-reminding agent runs are managed in one place, along with the creator versus agent identity choice for unattended tasks.