Skip to contentSkip to stories

Updated

#Tutorial/How-to

Items with an AI score under 20 are hidden. Show low-relevance items

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

Apr 17

Apr 17Fri

Apr 16

Apr 16Thu

Apr 6

Apr 6Mon
  1. Z.ai Release NotesOfficialAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

Apr 4

Apr 4Sat
  1. Andrej KarpathyXAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. Andrej KarpathyXAI score49

    Karpathy shares an LLM-maintained personal knowledge base workflow

    AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.

Mar 24

Mar 24Tue
  1. Anthropic EngineeringOfficialAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringOfficialAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

Mar 11

Mar 11Wed
  1. Nano Banana 2.1OfficialAI score62

    How to get the most out of Nano Banana 2 for image generation

    AINano Banana 2, also called Gemini 3.1 Flash Image, adds visual grounding with Google Search, 512px resolutions, and extreme aspect ratios of 1:8 and 1:4. The guide advises using it as the default for new projects, with Nano Banana Pro reserved for complex prompts it fails, and keeping Thinking mode off by default.

    Why it matters: The guide compares Nano Banana 1, 2, and Pro with concrete routing advice, which helps developers decide which model to default to and how to control cost.

Mar 10

Mar 10Tue
  1. Nick TurleyXAI score32

    ChatGPT adds interactive visual explanations for math and science concepts

    AIOpenAI rolled out interactive visual explanations in ChatGPT that let users adjust variables and manipulate formulas to see changes instantly in graphs and outcomes. The team built a Codex workflow that helps convert common math questions into visual learning blocks.

    Image from @nickaturley's post

Mar 2

Mar 2Mon
  1. Nano Banana 2.1OfficialAI score23

    Nano Banana 2 adds precise details through short prompt sentences

    AIGoogle's Nano Banana 2 image model lets users add exact details to outputs by writing short sentences in prompts. The post contrasts a bare prompt for a snow leopard portrait with a detailed one specifying pose, melting snow, flowers, a sun dog, lighting, and blue eyes.

    Image from @NanoBanana's post
  2. Hamel HusainBlogAI score44

    Hamel Husain and Shreya Shankar Release Evals Skills for Coding Agents

    AIHamel Husain and Shreya Shankar published evals skills, a set of skills for AI product evals that helps users avoid common mistakes. The entry point evals-start routes users to eval-audit for existing pipelines or error-discovery for unanalyzed traces. The repository is available at ai-evals-course/evals-skills and installs via npx skills add.

Mar 1

Mar 1Sun
  1. Artificial IgnoranceBlogAI score46

    Build Your Own Benchmark: Why Public AI Evals Are Saturating and What Replaces Them

    AIPublic AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro. OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory. The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.

Feb 27

Feb 27Fri

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 22

Feb 22Sun
  1. Artificial IgnoranceBlogAI score62

    Harness engineering emerges as a playbook for managing coding agents

    AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 19

Feb 19Thu

Feb 4

Feb 4Wed
  1. Anthropic EngineeringOfficialAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Jan 20

Jan 20Tue
  1. Anthropic EngineeringOfficialAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

Jan 13

Jan 13Tue
  1. Tim DettmersBlogAI score36

    Tim Dettmers shares a guide to automating your own work with coding agents

    AITim Dettmers, a professor who has used Claude Code for eight months, says more than 90% of code and text should be written by agents. He says most tasks cannot benefit from parallel agent sessions the way software engineering does. The post is a personal guide to automating one's own work, drawing on his experience beyond coding, including writing blog posts, grant proposals and meta reviews.

Dec 12, 2025

Dec 12, 2025Fri
  1. Nano Banana 2.1OfficialAI score24

    Nano Banana shares a prompt for cute isometric diorama lamps

    AINano Banana, the account of Google's Gemini team, shares a prompt template for generating cute isometric 3D cube diorama lamps. The prompt specifies internal lighting, chibi figurine styling, matte PVC material, a neutral background in a dark room, many small details, and subtle dust and scratch textures.

    Image from @NanoBanana's post

Dec 10, 2025

Dec 10, 2025Wed
  1. Andrej KarpathyBlogAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Nov 18, 2025

Nov 18, 2025Tue
  1. Quoc LeXAI score34

    Gemini 3 Deep Think builds an interactive analog film photography demo

    AIQuoc Le prompted Gemini 3 Deep Think to create a self-contained HTML analog film photography experience, and it produced a demo. The prompt required a viewfinder for uploaded images, adjustable settings, selectable film stocks that affect photo style, and a gallery of overlapping Polaroid-style prints that enlarge on hover.

    Video from @quocleix's post

Oct 29, 2025

Oct 29, 2025Wed

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

  2. Mira MuratiXAI score54

    Thinking Machines explores on-policy distillation for training small models

    AIThinking Machines published a post on on-policy distillation, a training approach combining the error-correcting relevance of RL with the reward density of SFT. The quoted post reports that in math reasoning and an internal chat assistant, on-policy distillation can outperform other approaches at a fraction of the cost.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabOfficialAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 27, 2025

Aug 27, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score23

    Cognition's Guide Shows How to Build Your Own AI Data Analyst

    AICognition published a guide titled "Build Your Own AI Data Analyst" on its blog, but the provided text contains only a navigation listing of unrelated posts and no guide content. No tools, models, features, or steps can be verified from the source.

Aug 4, 2025

Aug 4, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score22

    Devin Can Automate Migrating Jenkins Pipelines to GitHub Actions at Scale

    AICognition says generative AI agents such as Devin can read internal docs and convert Jenkins pipelines to GitHub Actions syntax, replacing custom plugins with Actions or APIs and validating the results. The company claims enterprises can cut multi-year migration efforts to a few months, with Devin running inside the customer's secure environment so code and secrets stay internal.

Jun 26, 2025

Jun 26, 2025Thu

Jun 22, 2025

Jun 22, 2025Sun
  1. Cognition Blog (Devin, Windsurf)OfficialAI score57

    Cognition details blockdiff, an open-source file format for instant VM disk snapshots

    AICognition built and open-sourced blockdiff, a file format that creates block-level diffs of VM disks using only filesystem metadata. The company reports its otterlink hypervisor cut snapshot times from 30 to 60 minutes on EC2 to about 5 to 10 seconds for a 128 GB disk with a 5 GB diff, roughly a 200x speedup. The post also covers the sparse file and copy-on-write concepts behind the approach and why OverlayFS, ZFS, and qcow2 were not chosen.

Feb 12, 2025

Feb 12, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score36

    Linktree Uses Devin to Add Social Platforms and Ship About 100 Merged PRs

    AILinktree has used Devin, Cognition's AI software engineer, to merge roughly 100 pull requests in a month, mostly fixing customer-reported bugs and implementing small features. The engineering team also used Devin to add support for new social media platforms, launching five Devins, one per repo and PR, and later used the Devin API with a Playbook script to spawn multiple Devins for multi-repo features. The team says Devin works best on tasks an engineer could finish in a couple of hours.

Jan 20, 2025

Jan 20, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score46

    Devin 101: Automatic PR Reviews with the Devin API

    AICognition's Devin can be triggered through its External API by GitHub Actions to automatically review pull requests, typically within five to ten minutes. The setup involves adding a workflow file, storing a DEVIN_API_KEY secret, and customizing the review prompt to match team conventions. Cognition recommends treating Devin as an extra reviewer rather than a replacement for human oversight, since it does not catch every bug.

Jan 13, 2025

Jan 13, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Crossmint Uses Devin to Scale Open-Source Development of GOAT SDK

    AICrossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.

Jul 19, 2024

Jul 19, 2024Fri
  1. Cognition Blog (Devin, Windsurf)OfficialAI score47

    Devin Automates Recovery of Windows Machines From CrowdStrike Outage Failure

    AICognition tested whether its Devin AI agent could recover Windows machines stuck in the Blue Screen of Death after the CrowdStrike outage. Following a playbook of eight steps, Devin mounted the drive, deleted the faulty CrowdStrike files, and debugged a failed volume detach. The blog says the machine was confirmed bootable, and that playbooks are most useful when the same fix must be repeated across many machines.

Jun 4, 2024

Jun 4, 2024Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score55

    Devin June 2024 update adds playbooks, snapshots, and Slack, Github, and Linear triggers

    AIDevin's June 2024 update lets users directly operate its VSCode editor, terminal, and browser, and adds Playbooks for recurring engineering tasks. It also adds Knowledge sharing, Machine Snapshots that persist dev environments, event-driven triggers from Slack, GitHub, and Linear, a Secrets manager, and tools for verifying Devin's work. Access remains limited to a waitlist with weekly invite releases.