Skip to contentSkip to stories

Updated

#Deployment/Engineering

Jul 8

Jul 8Wed
  1. Cognition Blog (Devin, Windsurf)AI score47

    Cognition Tests Trustworthiness of SWE-1.7, Built on Kimi K2.7 Code

    AICognition says its SWE-1.7 model, developed from the open-source Kimi K2.7 Code base, performs as well as or better than leading U.S. frontier models on its new trustworthiness evaluation suite. The suite combines 145 politically sensitive questions, sampled in English and Chinese, with realistic coding scenarios to measure propaganda, censorship, and security behavior. Cognition says SWE-1.7 improves substantially over the base Kimi K2.7 Code model, though the company says the benchmarks are still in development.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

Jul 6

Jul 6Mon

Jul 5

Jul 5Sun

Jul 3

Jul 3Fri
  1. Lil'Log (Lilian Weng)AI score62

    Lilian Weng surveys harness engineering as a path to recursive self-improvement

    AIThe post argues that the system surrounding a base model, called the harness, increasingly determines how well AI agents deploy and improve. It reviews research where harness components such as workflows, context, and code are optimized automatically through evolutionary search and meta-agent loops. The author concludes that evaluators, memory management, and human oversight remain open bottlenecks.

  2. Arthur MenschAI score34

    Mistral urges enterprises to adopt open-source models and own their AI data

    AIMistral CEO Arthur Mensch argues enterprises should use open-source models and store their own data to avoid dependence on closed providers that retain customer data. He says companies should build continuous training loops from employee and user interactions, and shrink models to cut deployment costs. Mistral positions its Studio control plane and Forge training platform, deployed on customer infrastructure or via zero-data-retention hosting, as tools for this shift.

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)AI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.

Jul 1

Jul 1Wed
  1. PromptArmor Threat IntelligenceAI score58

    Copilot Cowork Skills Still Reach DeepSeek After Admin Opt-Out

    AIPromptArmor reports that Skills in Microsoft Copilot Cowork can call DeepSeek even when an organization has not opted into the DeepSeek Preview. The calls use the agent's own access path, so users need no API key, and a Skill built this way received a 100/100 score from Microsoft's Skill Scanner. After Microsoft removed the DeepSeek Preview setting on June 25, the report says admins had no remaining setting to block DeepSeek through the Cowork code environment, leaving disabling Cowork entirely as the only option.

Jun 30

Jun 30Tue
  1. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

  2. Werner VogelsAI score22

    Werner Vogels says two-pizza teams are about ownership, not food

    AIAmazon CTO Werner Vogels argues that the "two-pizza" team concept was never about feeding engineers but about ownership, speed, and avoiding bureaucracy. He says working backwards from the customer and writing documents to force clarity remain core practices. He adds that the industry is changing and it is time to reconsider how products are brought to life.

  3. Tri DaoAI score53

    Tri Dao Praises Etched's Fast Inference Chip Design for LLM Serving

    AITri Dao says Etched designed and produced its chips within two years by hardcoding attention into silicon and reaching high MFU. He expects hardware built for LLM inference to cut the cost of intelligence by 10x. The quoted Etched post says it has built its first racks after an A0 tapeout, raised $800m, holds $1B+ in customer contracts, and plans to ship the racks this summer.

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

  2. Hamel HusainAI score54

    Why Hard-to-Eval AI Products Need Designs That Support Verification

    AIHamel Husain argues that an AI product whose output is hard to verify is a product design problem, not just an evaluation problem. He shows before-and-after sketches for an AI data agent, a PE lesson planner, and a workers' compensation report tool, each adding provenance, scoped edits, and checkable evidence. He notes that designing for verification also makes evals easier to build and grade.

Jun 28

Jun 28Sun
  1. PaddlePaddleAI score46

    PaddlePaddle announces Unlimited-OCR now runs in vLLM

    AIUnlimited-OCR, Baidu's long-context OCR model, now runs in vLLM, with a recipe provided for developers to try it. The background post says it parses entire books in one pass using Reference Sliding Window Attention (R-SWA), which keeps the KV cache fixed during decoding, and claims 35% faster throughput than DeepSeek-OCR at 6K output tokens.

Jun 27

Jun 27Sat
  1. PaddlePaddleAI score36

    PaddleFormers 1.2 adds DeepSeek-V4 training with 128K+ context support

    AIPaddleFormers 1.2 is released with support for training DeepSeek-V4 and 128K+ long-context training. The update adds Context Parallel, Packing, Document Mask Attention, and the Muon optimizer, plus ultra-fused mHC, CSA, and HCA operators, DeepEP/HybridEP communication, and lossless FP8 training with AutoSubbatch memory balancing. The project is presented as fully open-source and is available on GitHub.

  2. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 26

Jun 26Fri
  1. PaddlePaddleAI score32

    PP-OCRv6 Ep.4 benchmarks show 3.9x CPU speedup and 0.13s A100 OCR

    AIPaddlePaddle's PP-OCRv6 Tech Deep Dive Ep.4 benchmarks the OCR models across A100, V100, Intel Xeon CPU, and Apple M4 setups. PP-OCRv6_tiny processes an image in 0.13s on A100, while PP-OCRv6_tiny with OpenVINO runs 3.9x faster than PP-OCRv5_mobile on Intel CPU. The post recommends Medium for high-concurrency APIs, Small for CPU document systems, Tiny for mobile or embedded devices, and Medium or Small for multilingual business use.

  2. HyperdimensionalAI score62

    Dean W. Ball proposes private audits and certification for frontier AI labs

    AIDean W. Ball argues that the current government restrictions on frontier model releases amount to a de facto preapproval regime without a known safety standard. He proposes that independent verification organizations audit labs against their own safety frameworks, with government certifying or licensing the auditors. The post also argues that broad distribution of frontier AI is needed to learn what good safety practice looks like.

Jun 25

Jun 25Thu
  1. Lilian WengAI score40

    Lilian Weng's Overview of Scaling Laws and Compute-Optimal Allocation

    AILilian Weng published a long blog post on scaling laws, which help estimate the best split of compute between data and model size before a large training run. The post covers what scaling laws predict, how compute-optimal allocation works, and why Kaplan et al. and Chinchilla reach different conclusions. It also addresses how data limits and fitting details make extrapolation difficult.

Jun 24

Jun 24Wed

Jun 23

Jun 23Tue

Jun 20

Jun 20Sat

Jun 19

Jun 19Fri

Jun 18

Jun 18Thu
  1. Andrew NgAI score15

    DeepLearning.AI launches course on adding voice to AI agents

    AIDeepLearning.AI has launched a course, taught by VocalBridge CEO Ashwyn, on adding voice to AI agents and applications. It covers building voice agents that are both reliable and fast, with three projects: a voice-interactive game, an agent that gains a voice in about 10 lines of code, and an agent that places outbound calls via a make_phone_call function.