Skip to contentSkip to stories
Updated

Coding

Oct 7

Oct 7Wed
  1. Epoch AIOfficialAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  2. Hugging Face BlogOfficialAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

Oct 5

Oct 5Mon
  1. GitHub Blog · AI & MLOfficialAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

Oct 4

Oct 4Sun
  1. Epoch AIOfficialAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

Sep 29

Sep 29Tue
  1. Replit BlogOfficialAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogOfficialAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Jul 28

Jul 28Tue
  1. JetBrains AI BlogOfficialAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 13

Jul 13Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

Jun 8

Jun 8Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score70

    Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

    AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

    Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

Jun 3

Jun 3Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition Estimates Engineering Hours Saved by Its Devin Coding Agent

    AICognition built an automated agent that classifies Devin sessions as productive and estimates the human engineering hours each one would have taken. On 233 held-out sessions the estimator reached an rlog of 0.74, with individual errors often 2 to 3 times in either direction but roughly unbiased in aggregate. The system is calibrated to underestimate and is currently running with Devin customers.

    Why it matters: The post shows how the measurement design, from hours-based metrics to conservative calibration, determines whether agent productivity estimates can be trusted in aggregate.

Apr 30

Apr 30Thu
  1. OpenAI Alignment Research BlogOfficialAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringOfficialAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

That’s everything