Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 26

Sep 26Sat
  1. Varun MohanXAI score23

    lol, get that we’re getting memed for this but a bit of context.

    AIWe added planning mode in 2025 and deleted it from the product earlier this year. Users wanted a way to explicitly plan with the model so we added this opt in slash command. Understood that the timing couldn’t be worse since it appears like we’re adding this for the first time. Have a great weekend folks, lots more to come in the coming weeks!

  2. Alexander DoriaXAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

Sep 25

Sep 25Fri
  1. Lydia Hallie ✨XAI score22

    Claude Code's prompt-audit command renamed from /claude-api to /checkup

    AIAnthropic's Lydia Hallie says the prompt-audit command is now also available as /checkup, replacing the API-specific name that suggested it only worked with the API. The command checks CLAUDE.md, skills, and agents for instructions the model no longer needs, and it has always worked on Claude Code setups.

    Video from @lydiahallie's post
  2. Lydia Hallie ✨XAI score38

    Claude Code now stops at a graceful point when hitting usage limits

    AIClaude Code will now look for a graceful stopping point when a user hits the 5-hour limit mid-task, rather than cutting off mid-edit. It draws a small, fixed allowance from the weekly limit to finish what it can. The update responds to a frequently requested change.

  3. GitHub Blog · AI & MLOfficialAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

  4. Noah ZwebenXAI score46

    Anthropic shows Claude Code /remote-control demo with Opus 5.5 claymation video

    AIAnthropic's Noah Zweben shared a claymation video showing Claude Code's /remote-control feature, now made with Opus 5.5 after an earlier Opus 4.6 version. The feature, which lets users control Claude Code remotely, is rolling out to Pro users at 10% and ramping, with Team and Enterprise support coming later.

    Video from @noahzweben's post
  5. Microsoft CopilotOfficialAI score40

    Microsoft Copilot app refreshed to unify chat, agents, app building, and workflows

    AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.

  6. Satya NadellaXAI score52

    Satya Nadella announces Copilot update with Autopilot, Code, Home, and Office

    AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.

    Video from @satyanadella's post
  7. François CholletXAI score32

    Chollet: Software engineering difficulty stays constant across abstraction levels

    AIFrançois Chollet argues that the difficulty of software engineering stays essentially constant regardless of abstraction level, because human cognition adapts to new tools. He says tools are affordances rather than magic wands that eliminate work, and that great software engineering remains immensely challenging despite changed workflows. Simon Willison's background post similarly argues that coding agents make software engineering harder, requiring extraordinary discipline and knowledge.

Sep 24

Sep 24Thu
  1. GitHub Blog · AI & MLOfficialAI score46

    GitHub Copilot app's canvases argue chat is the wrong AI interface

    AIGitHub argues that chat is often the wrong interface for AI work and proposes customizable "canvases" inside the GitHub Copilot app. Canvases are full-stack applications running without browser chrome that can communicate bi-directionally with the Copilot agent and execute code locally. The post cites examples including a Connect 4 game, a Winget package manager UI, and a SQLite database interface.

  2. Lewis Tunstall @ COLM 🌉XAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  3. GitHub Blog · AI & MLOfficialAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

Sep 23

Sep 23Wed
  1. Amp NewsOfficialAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  2. Amp NewsOfficialAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

  3. Boris ChernyXAI score30

    Claude models tricky code states to find and fix bugs

    AIClaude builds a model of a program's most complex parts, such as state machines or race-prone code, and searches that model for counterexamples that signal suspected bugs. It then reproduces those bugs and fixes them in the code. The post clarifies that the whole codebase is not formally verified, only the riskiest sections are modeled and checked.

  4. Julien ChaumondXAI score36

    Hugging Face releases JS package to browse LeRobot datasets directly

    AIHugging Face has released a new JavaScript package, huggingface/lerobot, that lets developers read LeRobot datasets on the Hub directly in the browser without downloading them. The post says coding agents can use it to build custom dataset viewers quickly, with an example UI implemented in about 600 lines of JS.

    Image from @julien_c's post
  5. Karl's AI WattsXAI score22

    Notch admits he is enjoying vibe coding after earlier opposing AI coding

    AIMinecraft creator Notch says on X that he is enjoying vibe coding and admits he may have been slightly wrong. Months earlier he had publicly rejected AI-written code, but he later began having AI build internal tools such as a map editor and node graph tools. The main post adds that he has accumulated a set of small tools for himself before much game development has happened.

  6. Mike KnoopXAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  2. Alex AlbertXAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  3. ChatGPTOfficialAI score62

    OpenAI rolls out GPT-6 Sol and GPT-6 Luna in ChatGPT Work and Codex

    AIOpenAI announced GPT-6 Sol and GPT-6 Luna, rolling out today in ChatGPT Work and Codex. The rollout covers Plus, Pro, Business, Enterprise, and Edu users.

    Why it matters: The post names two new GPT-6 variants and their rollout to specific ChatGPT and Codex plan tiers, which shows how access is being staged.

    Video from @ChatGPT's post
  4. Mike KriegerXAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  5. Boris ChernyXAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

    Why it matters: The author reports a same-task comparison in which Opus 5.5 finished a HAProxy C-to-Rust port faster and cheaper than Fable 5.1, offering a concrete cost and time benchmark.

  6. catXAI score62

    Claude Opus 5.5 becomes the default model in Claude Code and Claude app

    AIClaude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. Anthropic is defaulting to effort medium across products, which it says is comparable to Fable 5.1 on intelligence but faster. Rate limits will go 25% further on Opus 5.5 compared to Opus 5.

    Why it matters: The post names concrete default changes across Claude Code and the Claude app, plus a specific effort setting and rate-limit difference, useful for judging day-to-day cost and speed.