Skip to contentSkip to stories

Updated

#Coding

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. 🚨 AI News | TestingCatalogAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    AIAnthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

    Image from @testingcatalog's post
  2. Claude Code · GitHub ReleasesAI score36

    Claude Code v2.1.293 adds Claude Haiku 5.5 and fixes dozens of bugs

    AIClaude Code v2.1.293 adds Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API, with 1M context and pricing of $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K). The release also adds agentType to the subagentStatusLine payload and isDeferred to $.tool.register, and fixes numerous issues including a memory leak in HTTP MCP connections.

  3. IThome · AIAI score42

    Nvidia unveils DGX Station for Windows, a desktop AI supercomputer for running trillion-parameter models

    AINvidia announced DGX Station for Windows, a desktop AI supercomputer built on the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip with up to 748GB of unified memory, able to run models of up to about one trillion parameters locally. The machine offers up to 20 PFLOPS of AI compute and combines 252GB of HBM3e GPU memory with 496GB of LPDDR5X CPU memory. It is scheduled to go on sale in the fourth quarter of 2026.

  4. Design ArenaAI score44

    Claude Opus 5.5 tops four Design Arena leaderboards after two weeks

    AIAnthropic's Claude Opus 5.5 has taken first place on four Design Arena leaderboards: Overall Frontend, Data Visualization, 3D Design, and React Native. It also ranks in the top three on the Game Dev and UI Components leaderboards, about two weeks after its release. Design Arena says developers, designers, and casual users have embraced the model.

    Image from @DesignArena's post
  5. GitHub Copilot ChangelogAI score30

    GitHub launches purpose-built AI model for leaked secret detection across developer workflows

    AIGitHub is rolling out a fine-tuned, purpose-built model for secret detection that reads surrounding code to identify likely credentials, including passwords without recognizable token formats. Existing AI-detected Password alerts have been upgraded automatically, and AI-detected secrets in push protection is in private preview. New opt-in checks in push protection and the GitHub Copilot /security-review command will consume GitHub AI Credits.

  6. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  7. The SequenceAI score37

    The Sequence Learning Loop: OpenAI DevDay and Gemini 4 Argon Show Workflow Competition

    AIThe newsletter argues that AI competition is shifting toward completed workflows, citing OpenAI's September 29 DevDay announcements on cost and infrastructure and Google's September 30 introduction of Gemini 4 Argon for longer, more demanding reasoning tasks. It says coding agents must inspect repositories, edit code, run tests, and deliver reviewable work, so cost, context, and supervision matter alongside model intelligence.

  8. Claude BlogAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

Oct 6

Oct 6Tue
  1. meng shaoAI score35

    Claude Code's html-plan plugin turns plans into reviewable HTML pages

    AIClaude Code developer Thariq (@trq212) released html-plan, a plugin that makes Claude Code generate self-contained single-file HTML plans instead of lengthy Markdown. The page organizes the plan into a layered tree with progressive disclosure, numbered decision points, and in-page feedback that can be pasted back into Claude Code. Install it with claude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install html-plan@claude-community.

    Image from @shao__meng's post
  2. meng shaoAI score48

    Independent review layer keeps LLM data agent from judging its own SQL

    AIA data analysis agent built by @Sumanth_077 separates generation, deterministic guardrails, and review: Qwen writes read-only SELECT queries, code enforces hard rules such as a single SELECT, SQLite read-only mode, and a 200-line limit, and a separate TypeSafe AI Jev model checks question clarity, SQL relevance, and whether answers are grounded in returned rows. Answers that fail grounding are marked as unverified drafts while the SQL and data are kept for human inspection.

    Image from @shao__meng's post
  3. meng shaoAI score30

    MIT 6.S950 Lecture 4 Explores Programming's Abstraction Ladder in the AI Era

    AIMIT's 6.S950 "Agency with AI" course has released Lecture 4, "The Abstraction Ladder (of Programming)," which compares today's prompt-driven coding with the 1957 FORTRAN paper by Backus et al. The lecture argues that the objections to vibe coding echo the arguments once raised against compilers, but natural-language "compilation" differs because the same prompt can yield different programs each time, unlike deterministic translation.

    Image from @shao__meng's post
  4. GitHub Copilot ChangelogAI score32

    Update your IDE to restore Copilot agent activity in usage metrics

    AIGitHub says some IDEs that moved Copilot agent sessions to the Copilot SDK left that activity unattributed in usage metrics, and a fix is rolling out by IDE. Visual Studio Code 1.139.0 and later has the fix now, while Visual Studio 18.12, JetBrains, Eclipse, and Xcode are expected between October and November 2026. Billing is unaffected, and missing data from affected versions cannot be backfilled.

  5. GitHubAI score72

    GitHub rebuilds Git infrastructure to handle agent-scale write volume

    AIGitHub reports that Git events on the platform rose from 218.2 billion to 473.3 billion per month between September 2025 and August 2026. It says agent workloads push write throughput and merge contention beyond what its current replica-based architecture handles well, so it is separating durable storage from compute while GitHub keeps running. The article states internal benchmarks reached up to 35 times higher write throughput.

    Why it matters: The post links rising Git event volume to specific architectural bottlenecks, showing why agent workloads strain write paths and how GitHub plans to separate storage from compute.

  6. Gemini CLI · GitHub ReleasesAI score14

    Gemini CLI v0.63.0 released with retry indicator and auth loop fixes

    AIGemini CLI v0.63.0 adds a retry progress indicator during connection recovery and fixes an infinite authentication loop caused by file contention, headless keyring issues, and supervisor state drops. The release also bounds tool output size and cleans up temporary directories when background shell execution exits, alongside fixes for MCP enablement config handling and stdin restoration after capability detection.

  7. Boris ChernyAI score38

    Boris Cherny shares prompts for formally verifying Claude Agent SDK

    AIBoris Cherny says he used Opus 5.5 with Lean to formally verify the Claude Agent SDK, with a couple of short prompts producing 16 PRs fixing bugs and race conditions. He also reports that TLA+ works well, sometimes combined with Lean to find data flow, concurrency, and state management issues. The post links to his actual prompts as another example.

  8. AnthropicAI score49

    Anthropic expands Cyber Verification Program for verified security professionals

    AIAnthropic is expanding its Cyber Verification Program to give verified security professionals broader access to its most capable models. Through the program, they can use Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work. New tiers will also allow authorized offensive work such as penetration testing and red-teaming.

  9. Claude Code · GitHub ReleasesAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    AIClaude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.