Skip to contentSkip to stories

Updated

#Coding

Showing low-relevance items too. Hide low-relevance items

Sep 26

Sep 26Sat
  1. KhazixXAI score18

    Khazix says Claude Opus 5.5 excels at long-running agent tasks

    AIThe author ran a full-rewrite-scale task with Claude Opus 5.5 for over six hours, reading every thinking summary along the way. They call it stable and strong, rating it top tier across communication, comprehension, aesthetics, development, and long-horizon agent work, and wish OpenAI would catch up.

    Image from @Khazix0918's post
  2. Alexander DoriaXAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

Sep 25

Sep 25Fri
  1. GitHub Blog · AI & MLOfficialAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

  2. Mustafa SuleymanXAI score14

    Microsoft's Copilot team launches Autopilot, per Mustafa Suleyman

    AIMustafa Suleyman praised the Copilot team's work and promoted a new feature called Autopilot, directing readers to check it out. The post provides no details about Autopilot's functionality, capabilities, or availability. Background from a linked article references Home, Code, and Autopilot sections but does not establish specifics.

  3. Kilo (acq. by Anaconda)OfficialAI score10

    Kilo argues engineers can skip this week's AI launches

    AIKilo's blog argues that most engineers can skip this week's AI launches without losing anything. It recommends testing tools against one slow part of your real work rather than tracking which model tops the leaderboard.

    Image from @kilocode's post
  4. Microsoft CopilotOfficialAI score40

    Microsoft Copilot app refreshed to unify chat, agents, app building, and workflows

    AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.

  5. Satya NadellaXAI score52

    Satya Nadella announces Copilot update with Autopilot, Code, Home, and Office

    AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.

    Video from @satyanadella's post
  6. François CholletXAI score32

    Chollet: Software engineering difficulty stays constant across abstraction levels

    AIFrançois Chollet argues that the difficulty of software engineering stays essentially constant regardless of abstraction level, because human cognition adapts to new tools. He says tools are affordances rather than magic wands that eliminate work, and that great software engineering remains immensely challenging despite changed workflows. Simon Willison's background post similarly argues that coding agents make software engineering harder, requiring extraordinary discipline and knowledge.

Sep 24

Sep 24Thu
  1. GitHub Blog · AI & MLOfficialAI score46

    GitHub Copilot app's canvases argue chat is the wrong AI interface

    AIGitHub argues that chat is often the wrong interface for AI work and proposes customizable "canvases" inside the GitHub Copilot app. Canvases are full-stack applications running without browser chrome that can communicate bi-directionally with the Copilot agent and execute code locally. The post cites examples including a Connect 4 game, a Winget package manager UI, and a SQLite database interface.

  2. Lewis Tunstall @ COLM 🌉XAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  3. GitHub Blog · AI & MLOfficialAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

Sep 23

Sep 23Wed
  1. Amp NewsOfficialAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  2. Amp NewsOfficialAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

  3. Boris ChernyXAI score30

    Claude models tricky code states to find and fix bugs

    AIClaude builds a model of a program's most complex parts, such as state machines or race-prone code, and searches that model for counterexamples that signal suspected bugs. It then reproduces those bugs and fixes them in the code. The post clarifies that the whole codebase is not formally verified, only the riskiest sections are modeled and checked.

  4. howie.seriousXAI score18

    Opus 5.5 at medium thinking produces a near-professional promo video

    AIHowie.serious says Opus 5.5 with only medium thinking enabled produced a video of impressive quality, suggesting a grim outlook for human video work. The quoted post says the model made a promo video for the Applore app in a few minutes, and the author calls the result remarkable.

    Video from @howie_serious's post
  5. Karl's AI WattsXAI score22

    Notch admits he is enjoying vibe coding after earlier opposing AI coding

    AIMinecraft creator Notch says on X that he is enjoying vibe coding and admits he may have been slightly wrong. Months earlier he had publicly rejected AI-written code, but he later began having AI build internal tools such as a map editor and node graph tools. The main post adds that he has accumulated a set of small tools for himself before much game development has happened.

  6. Mike KnoopXAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post