Skip to content

#Tutorial/How-to

Feb 27

Feb 27Fri
  1. Max WoolfAI score38

    New blog post up: the culmination of my past few months working with agents Opus 4.5 and beyond, and the *many* things I learned. Also, the discovery of an optimization trick with promise. As a bonus: this post will make Rust engineers very mad. https://minimaxir.com/2026/02/ai-agent-coding/

    New blog post up: the culmination of my past few months working with agents Opus 4.5 and beyond, and the *many* things I learned. Also, the discovery of an optimization trick with promise. As a bonus: this post will make Rust engineers very mad. https://minimaxir.com/2026/02/ai-agent-coding/

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)AI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    Cognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    AIWhy it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 22

Feb 22Sun
  1. Artificial IgnoranceAI score62

    Harness engineering emerges as a playbook for managing coding agents

    The article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 19

Feb 19Thu
  1. Eugene YanAI score27

    yay for automatic prefix caching! just set a single cache control field at the top level of your request body. note that you do still need to structure your prompt templates to benefit from prefix caching though. https://platform.claude.com/docs/en/build-with-claude/prompt-caching#automatic-caching

    yay for automatic prefix caching! just set a single cache control field at the top level of your request body. note that you do still need to structure your prompt templates to benefit from prefix caching though. https://platform.claude.com/docs/en/build-with-claude/prompt-caching#automatic-caching

Feb 6

Feb 6Fri

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    Nicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    AIWhy it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    Ai2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Jan 20

Jan 20Tue
  1. Anthropic EngineeringAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    Anthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    AIWhy it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

Jan 13

Jan 13Tue
  1. Tim DettmersAI score36

    Tim Dettmers Argues Agents Should Automate Most Personal Work, Not Just Code

    Tim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.

Dec 10, 2025

Dec 10, 2025Wed
  1. Andrej KarpathyAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    Andrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Nov 18, 2025

Nov 18, 2025Tue

Oct 29, 2025

Oct 29, 2025Wed

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    Cognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    Thinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    AIWhy it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    Cognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    AIWhy it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    Eight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 27, 2025

Aug 27, 2025Wed

Aug 4, 2025

Aug 4, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score22

    Devin Can Automate Migrating Jenkins Pipelines to GitHub Actions at Scale

    Cognition says generative AI agents such as Devin can read internal docs and convert Jenkins pipelines to GitHub Actions syntax, replacing custom plugins with Actions or APIs and validating the results. The company claims enterprises can cut multi-year migration efforts to a few months, with Devin running inside the customer's secure environment so code and secrets stay internal.

Jun 26, 2025

Jun 26, 2025Thu

Jun 22, 2025

Jun 22, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score57

    Cognition details blockdiff, an open-source file format for instant VM disk snapshots

    Cognition built and open-sourced blockdiff, a file format that creates block-level diffs of VM disks using only filesystem metadata. The company reports its otterlink hypervisor cut snapshot times from 30 to 60 minutes on EC2 to about 5 to 10 seconds for a 128 GB disk with a 5 GB diff, roughly a 200x speedup. The post also covers the sparse file and copy-on-write concepts behind the approach and why OverlayFS, ZFS, and qcow2 were not chosen.

Feb 12, 2025

Feb 12, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score36

    Linktree Uses Devin to Add Social Platforms and Ship About 100 Merged PRs

    Linktree has used Devin, Cognition's AI software engineer, to merge roughly 100 pull requests in a month, mostly fixing customer-reported bugs and implementing small features. The engineering team also used Devin to add support for new social media platforms, launching five Devins, one per repo and PR, and later used the Devin API with a Playbook script to spawn multiple Devins for multi-repo features. The team says Devin works best on tasks an engineer could finish in a couple of hours.

Jan 20, 2025

Jan 20, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 101: Automatic PR Reviews with the Devin API

    Cognition's Devin can be triggered through its External API by GitHub Actions to automatically review pull requests, typically within five to ten minutes. The setup involves adding a workflow file, storing a DEVIN_API_KEY secret, and customizing the review prompt to match team conventions. Cognition recommends treating Devin as an extra reviewer rather than a replacement for human oversight, since it does not catch every bug.

Jan 13, 2025

Jan 13, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Crossmint Uses Devin to Scale Open-Source Development of GOAT SDK

    Crossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.

Jul 19, 2024

Jul 19, 2024Fri
  1. Cognition Blog (Devin, Windsurf)AI score47

    Devin Automates Recovery of Windows Machines From CrowdStrike Outage Failure

    Cognition tested whether its Devin AI agent could recover Windows machines stuck in the Blue Screen of Death after the CrowdStrike outage. Following a playbook of eight steps, Devin mounted the drive, deleted the faulty CrowdStrike files, and debugged a failed volume detach. The blog says the machine was confirmed bootable, and that playbooks are most useful when the same fix must be repeated across many machines.

Jun 4, 2024

Jun 4, 2024Tue
  1. Cognition Blog (Devin, Windsurf)AI score55

    Devin June 2024 update adds playbooks, snapshots, and Slack, Github, and Linear triggers

    Devin's June 2024 update lets users directly operate its VSCode editor, terminal, and browser, and adds Playbooks for recurring engineering tasks. It also adds Knowledge sharing, Machine Snapshots that persist dev environments, event-driven triggers from Slack, GitHub, and Linear, a Secrets manager, and tools for verifying Devin's work. Access remains limited to a waitlist with weekly invite releases.