Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 2

Oct 2Fri
  1. Design ArenaOfficialAI score42

    Astra adds accessibility code to games without being asked

    AIDesign Arena reports that every Astra-generated game in its code sample included accessibility features such as screen reader labels and reduced motion. In 60 of those games, the model never mentioned accessibility in its reasoning, suggesting it added these features by default.

  2. ClaudeDevsOfficialAI score29

    Anthropic adds "You should Know" plugin to Claude Code

    AIAnthropic is adding a new Claude Code plugin called "You should Know" that scans Claude's output for important information users might otherwise miss. It can be enabled with the command /plugin enable cc-plugin-you-should-know@builtin.

    Video from @ClaudeDevs's post
  3. Claude Code · GitHub ReleasesOfficialAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  4. SGLangOfficialAI score58

    SGLang v0.5.21 adds native decisions API and new model support

    AISGLang has released v0.5.21 with a native Decisions API that turns an LLM or VLM into a low-latency classifier and scorer. The release also lets /v1/score rerank search or RAG results in one call, lets PD instances switch between prefill and decode without restarting, and adds support for models including DeepSeek-V4.1 Flash, Kimi K3, and GLM-5.3-Flash on AMD MI355X. The announcement reports a 22% faster first token on long prompts for DeepSeek-V4.1 Flash and 20.6% higher prefill throughput for Kimi K3 in PD serving.

    Image from @sgl_project's post
  5. GitHub Copilot ChangelogOfficialAI score34

    Copilot code review gains API access and Balanced default effort level

    AIGitHub Copilot code review can now be requested through the REST and GraphQL APIs, with an optional review effort level set per request. Balanced became the default review effort level for new and existing repositories and organizations as of September 28, 2026, while users who explicitly selected Lite keep that setting. The changes are generally available to Copilot Pro, Pro+, Max, Business, and Enterprise plans.

  6. DatabricksOfficialAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  7. CursorOfficialAI score42

    Cursor's Rollouts detects deployment regressions and launches cloud agent fixes

    AICursor introduced Rollouts, a tool that writes a monitoring plan and watches changes as they deploy to catch regressions before users see them. When Rollouts detects a regression, it identifies the offending PR and opens an issue, and one click starts a cloud agent to fix it. Rollouts usage credits are included through Oct 3.

  8. GitHubOfficialAI score44

    GitHub Copilot adds Project HydraFusion and new models to model picker

    AIGitHub has made the Project HydraFusion research preview available in the GitHub Copilot app and @code, where it orchestrates multiple models rather than acting as a single model. New models from Anthropic (Fable 5.1 and Opus 5.5) and OpenAI (GPT-6.1 Sol) are also now selectable in the Copilot model picker.

  9. Microsoft CopilotOfficialAI score20

    Copilot Code lets more people build apps and workflows

    AIMicrosoft's Copilot Code is designed to help more people turn ideas into apps, workflows, and solutions for their work. Microsoft Copilot EVP Jacob Andreou discusses how the product expands who gets to build.

    Video from @MSFTCopilot's post
  10. Kilo (acq. by Anaconda)OfficialAI score36

    Ling 3.1 Flash is free in Kilo Code until October 13

    AIKilo Code is offering Ling 3.1 Flash for free until October 13, with the model served by Novita Labs. Ant Ling's background post describes the model as roughly 560B total parameters with about 25B active per token and a context window of up to 1M tokens. Ant Ling says it scores 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, and plans to open-source it soon.

  11. GitHub Blog · AI & MLOfficialAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.

  12. Latent SpaceBlogAI score43

    Airbnb CTO Ahmad Al-Dahle Details AI-Native Overhaul of Airbnb's Products and Workflows

    AIAirbnb CTO Ahmad Al-Dahle, who joined from Meta in January, says 60% of the company's code is now AI-authored and pull-request throughput per engineer is up about 1.6x. Roughly half of Airbnb's support tickets are now resolved purely by AI, which the company tested with synthetic data before production. Airbnb's internal context graph Everest helped speed up the grocery delivery and airport pickup services, which took eight to nine months and about six weeks to build, respectively.

  13. GitHub Copilot ChangelogOfficialAI score53

    GitHub Copilot adds new models, dynamic workflows, and desktop app automation

    AIGitHub Copilot's weekly release adds Claude Sonnet 5.5 and GPT-6.1 Sol for specified plan tiers, plus HydraFusion, a research preview that lets Copilot select and coordinate models for a task. It also introduces dynamic workflows in public preview, which let users save and reuse multi-step processes, and computer use in public preview on macOS and Windows for automating desktop apps.

  14. O'Reilly RadarBlogAI score39

    Coding Agents Benefit From Architectural Decision Records, With Limits

    AIArchitectural Decision Records (ADRs) give coding agents durable project context, helping them distinguish intentional decisions from implementation details. Agents can over-apply accepted but obsolete ADRs, so the author recommends explicit AGENTS.md instructions treating accepted ADRs as binding, prompting agents to flag conflicts, and keeping each ADR current rather than recording amendment logs.

  15. Lucas Beyer (bl16)XAI score45

    Lucas Beyer praises new coding benchmark for finding bugs in repos

    AILucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes. He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.

Oct 1

Oct 1Thu
  1. Latent.SpaceXAI score60

    Recursive Language Models explained by MIT's Alex Zhang on coding agents

    AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.

    Video from @latentspacepod's post
  2. NVIDIA BlogOfficialAI score62

    NVIDIA Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast mode in API

    AIGPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available in the OpenAI API and to eligible ChatGPT Work and Codex users. The source says Ultrafast offers up to 8x faster token generation than Astra Standard mode, which can shorten coding agents' response times between tool calls. OpenAI also says it uses its own models to keep optimizing inference software on NVIDIA GPUs after deployment.

    Why it matters: The source ties a specific speed claim to coding agents' edit-test-debug loops, showing where faster token generation changes developer workflows.

  3. ClineOfficialAI score34

    Cline reports DeepSeek V4 Pro costs about 30x less than Claude Opus 5

    AICline says two of its largest tasks in the last 30 days each processed 9B tokens, costing about $8,500 on Claude Opus 5 versus about $300 on DeepSeek V4 Pro. The post says that is roughly 30x cheaper for the same token count, letting users run large-horizon work without spending thousands or waiting on limit resets.

  4. Boris ChernyXAI score54

    Claude Code adds mods that customize behavior and UI via plugins

    AIClaude Code now supports mods that change its behavior, customize the UI, or add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app, and the author says mods can be shared so others can try them.

  5. Latent.SpaceXAI score55

    Anthropic's Thariq explains what Claude Code mods can access and control

    AIAnthropic's Thariq describes Claude Code mods, which can read conversation scope such as turn count and token usage. Mods run in process, so they can spawn subagents, parse their results with structured output, and modify the UI, which hooks cannot do. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app.

    Video from @latentspacepod's post
  6. Lydia Hallie ✨XAI score62

    Claude Code adds mods that customize behavior and UI via TypeScript plugins

    AIClaude Code can now be modified with mods that change its behavior, customize the UI, and add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app. A TypeScript function can intercept internal events such as tool calls, prompts, model requests, and renders, and add custom UI and commands.

  7. Comfy BlogOfficialAI score47

    Hakoniwa uses Comfy Agent to make the animated short YUI

    AIArtist 852 Hakoniwa made YUI, described as the first animated short created with Comfy Agent, which the ComfyUI team says took three days of focused work by one person at about 200,000 yen in total cost, excluding labor. The source says the film was made mostly with Seedance 2.5 and the making-of video with MiniMax H3, with Comfy Agent used to regenerate shots and compare video models.

  8. DatabricksOfficialAI score20

    Databricks Smart Routing assigns each coding task to a suitable model

    AIDatabricks' Smart Routing evaluates each coding task separately and selects the lowest-cost model capable of handling it, balancing quality, latency, and cost. In a demo, Omnigent splits an app build into planning, backend, and frontend work, routes each part to a different model, and runs some tasks in parallel.

    Video from @databricks's post
  9. Josh WoodwardXAI score34

    Google launches Stitch CLI to generate design ideas from terminal

    AIGoogle has introduced the @google/stitch CLI, letting users generate screens and design systems without leaving the terminal. It connects to local coding agents and can send a local dev server snapshot to Stitch. The tool complements the existing Stitch MCP and SDK, and can also be driven through agents such as Antigravity.

  10. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  11. OpenRouter · New modelsBlogAI score36

    Pareto 26.10 Preview: A Multimodal Model for Research, Coding and Agents

    AIPareto 26.10 Preview is a multimodal composite model built for research, coding, and agentic workflows. It is described as delivering frontier-level performance across a broad range of general-purpose tasks, though the source excerpt is a preview and provides no benchmark scores, parameter counts, pricing, or availability details.

  12. JetBrains AI BlogOfficialAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  13. Anthropic NewsroomOfficialAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.