Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Claude BlogOfficialAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

Oct 6

Oct 6Tue
  1. meng shaoXAI score35

    Claude Code's html-plan plugin turns plans into reviewable HTML pages

    AIClaude Code developer Thariq (@trq212) released html-plan, a plugin that makes Claude Code generate self-contained single-file HTML plans instead of lengthy Markdown. The page organizes the plan into a layered tree with progressive disclosure, numbered decision points, and in-page feedback that can be pasted back into Claude Code. Install it with claude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install html-plan@claude-community.

    Image from @shao__meng's post
  2. meng shaoXAI score48

    Independent review layer keeps LLM data agent from judging its own SQL

    AIA data analysis agent built by @Sumanth_077 separates generation, deterministic guardrails, and review: Qwen writes read-only SELECT queries, code enforces hard rules such as a single SELECT, SQLite read-only mode, and a 200-line limit, and a separate TypeSafe AI Jev model checks question clarity, SQL relevance, and whether answers are grounded in returned rows. Answers that fail grounding are marked as unverified drafts while the SQL and data are kept for human inspection.

    Image from @shao__meng's post
  3. meng shaoXAI score30

    MIT 6.S950 Lecture 4 Explores Programming's Abstraction Ladder in the AI Era

    AIMIT's 6.S950 "Agency with AI" course has released Lecture 4, "The Abstraction Ladder (of Programming)," which compares today's prompt-driven coding with the 1957 FORTRAN paper by Backus et al. The lecture argues that the objections to vibe coding echo the arguments once raised against compilers, but natural-language "compilation" differs because the same prompt can yield different programs each time, unlike deterministic translation.

    Image from @shao__meng's post
  4. Abida JuleXAI score22

    Top 10 Hermes agent skills ranked by GitHub stars on Reddit

    AIA Reddit thread prompted a ranking of the top 10 Hermes skills by GitHub stars, with the list spanning coding, knowledge graphs, and research tools. The entries include superpowers, an agentic skills framework that the post says works for software development, and a caveman-style skill and proxy that the post says cuts 65% of tokens for coding agents. Other listed items include a skill that researches topics across Reddit, X, YouTube, HN, Polymarket, and the web, and K-Dense-AI's collection of 165 validated scientific skills.

    Image from @I_am_Aiabir's post
  5. GitHub Copilot ChangelogOfficialAI score32

    Update your IDE to restore Copilot agent activity in usage metrics

    AIGitHub says some IDEs that moved Copilot agent sessions to the Copilot SDK left that activity unattributed in usage metrics, and a fix is rolling out by IDE. Visual Studio Code 1.139.0 and later has the fix now, while Visual Studio 18.12, JetBrains, Eclipse, and Xcode are expected between October and November 2026. Billing is unaffected, and missing data from affected versions cannot be backfilled.

  6. GitHubOfficialAI score72

    GitHub rebuilds Git infrastructure to handle agent-scale write volume

    AIGitHub reports that Git events on the platform rose from 218.2 billion to 473.3 billion per month between September 2025 and August 2026. It says agent workloads push write throughput and merge contention beyond what its current replica-based architecture handles well, so it is separating durable storage from compute while GitHub keeps running. The article states internal benchmarks reached up to 35 times higher write throughput.

    Why it matters: The post links rising Git event volume to specific architectural bottlenecks, showing why agent workloads strain write paths and how GitHub plans to separate storage from compute.

  7. Boris ChernyXAI score38

    Boris Cherny shares prompts for formally verifying Claude Agent SDK

    AIBoris Cherny says he used Opus 5.5 with Lean to formally verify the Claude Agent SDK, with a couple of short prompts producing 16 PRs fixing bugs and race conditions. He also reports that TLA+ works well, sometimes combined with Lean to find data flow, concurrency, and state management issues. The post links to his actual prompts as another example.

  8. AnthropicOfficialAI score49

    Anthropic expands Cyber Verification Program for verified security professionals

    AIAnthropic is expanding its Cyber Verification Program to give verified security professionals broader access to its most capable models. Through the program, they can use Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work. New tiers will also allow authorized offensive work such as penetration testing and red-teaming.

  9. Claude Code · GitHub ReleasesOfficialAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    AIClaude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.

  10. ClaudeDevsOfficialAI score20

    Claude Pro and Max users can claim one-time cloud sessions bonus credit

    AIAnthropic's ClaudeDevs account says Pro or Max plan subscribers on September 23 can still claim a one-time bonus credit for cloud sessions by running /claim-credit in Claude Code by October 7 at 11:59pm PT. Cloud sessions draw on this credit first before counting toward plan limits, and the credit expires November 4.

  11. ClaudeDevsOfficialAI score38

    Claude Code cloud sessions run parallel tasks on fresh VMs

    AIAnthropic's ClaudeDevs says Claude Code cloud sessions run each task on a fresh VM, so users can start several at once. The sessions keep running after the user closes their laptop. A field guide covers seven suitable workflows and how to connect GitHub.

  12. Mistral AIOfficialAI score47

    Mistral Large 4 solves 18 of 19 CTF challenges in speedrun test

    AIMistral Large 4 solved 18 of 19 challenges in a CTF speedrun, with tool calls and solve times drawn from actual runs. The post frames the model as efficient at reasoning over diverse complex challenges compared with other models.

    Video from @MistralAI's post
  13. laurenXAI score42

    Developer automates releases and QA with Grok Bot agents in Slack

    AIA developer used Grok Bot to build two Slack team bots, sandcastle for release management and poteto for engineering, automating their release and QA process. Sandcastle DMs contributors PR links, kicks off builds, and runs a fuzz swarm of 10+ Grok 4.7 xhigh agents, while poteto triages and fixes issues via Cursor Projects.

    Image from @poteto's post
  14. LangChainOfficialAI score46

    LangChain video shows how to build a model router into an agent harness

    AILangChain's Sydney Runkle presents a four-step method for building a model router into a coding agent harness: understanding tasks, understanding models, building the router, and tracking task outcomes. The background post says the router cut costs by 64% without reducing quality by sending tasks that do not need a frontier model to cheaper models.

  15. TiboXAI score23

    Codex adds "Approve for me" auto-review permission mode

    AICodex now offers an "Approve for me" permission mode, which automatically reviews actions instead of requiring manual approval. To enable it, open the permissions menu below the composer and select "Approve for me."

  16. Google LabsOfficialAI score57

    Google Flow Music Spaces can now export custom tools as VST3/AU plugins

    AIGoogle Flow Music lets creators build custom instruments or effects from natural language, and Spaces can now be exported as VST3/AU plugins. These plugins run inside producers' Digital Audio Workstations, so tools can fit existing production workflows. The source gives producer Khris Riddick-Tynes's "No Chaser" plugin as an example for checking instrumentals and vocals.

  17. Vaibhav (VB) SrivastavXAI score34

    Codex Auto-review now free for ChatGPT sign-in users

    AIOpenAI's Codex "Approve for me" mode uses a separate Auto-review agent to check actions needing approval, such as running commands outside the sandbox or accessing extra files and network resources. It reduces approval prompts during long tasks while keeping sandbox protections, and it is now free with ChatGPT sign-in without drawing from plan usage.

    Image from @reach_vb's post
  18. OpenRouter · New modelsBlogAI score62

    Mistral Large 4 is listed on OpenRouter with a 1M-token context window

    AIMistral AI's Mistral Large 4 is listed on OpenRouter as a frontier multimodal model accepting text and image input. The listing says it is built for reasoning, coding, and agentic workloads and offers a 1M-token context window. The feed excerpt is truncated, so further details such as pricing or availability are not confirmed here.

  19. Kilo (acq. by Anaconda)OfficialAI score29

    Kilo launches Kilo Desktop, a unified app for 500+ AI models

    AIKilo has launched Kilo Desktop, a single app offering access to more than 500 models from major labs, including open-source and local models. It includes agents that plan, code, and debug alongside users, plus built-in notebooks, local model support, and conda environments.

    Image from @kilocode's post
  20. Guillaume Lample @ NeurIPS 2024XAI score26

    Mistral model beats GLM 5.3 on STEM, CAD, and finance tasks

    AIOn human evaluation, the model outperforms GLM 5.3 on STEM, CAD, and finance tasks and performs on par on agentic coding. The post is part 5 of a thread, so the model's name and other details come from earlier posts not included here.

    Image from @GuillaumeLample's post
  21. Guillaume Lample @ NeurIPS 2024XAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  22. Gergely OroszXAI score36

    Uber uses AI to migrate 600,000 JUnit 4 tests to JUnit 5

    AIUber's engineers describe migrating 600,000 JUnit 4 tests covering 15 million lines of code to JUnit 5, which the post author says was impractical by manual means. The author says AI now makes such a large migration feasible, and points readers to Uber's engineering blog for the details.

    Image from @GergelyOrosz's post
  23. The SequenceBlogAI score62

    Darwin Gödel Machine rewrote its own scaffolding to raise SWE-bench scores

    AIThe Darwin Gödel Machine, a coding agent from Sakana and Jeff Clune's lab, modified its own codebase over roughly eighty iterations without supervision. Its additions included better file viewing, patch validation before submitting fixes, generating and ranking several candidate solutions, and keeping a history of failed attempts. These changes raised its score from 20 to 50 percent on SWE-bench and from 14 to 31 percent on Polyglot.

  24. Vaibhav (VB) SrivastavXAI score43

    Auto-review in Codex is now free for ChatGPT-signed-in users

    AIOpenAI has made Auto-review free for all users signed in through a ChatGPT account, and it does not draw usage from their plan. Auto-review uses a second agent to check the primary agent's actions, blocking high-risk moves and actions that drift from user intent, so long tasks can run without constant approval prompts. It can be enabled under settings > permissions > auto-review.

  25. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.