Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 6

Oct 6Tue
  1. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

  2. Harrison ChaseXAI score20

    Harrison Chase praises a take on agent harnesses

    AIHarrison Chase, founder of LangChain, endorsed a post on harnesses with the brief comment "Good take on harnesses." The post, from @zeeg, argues that general coding harnesses like Codex will be superseded by specialized ones and that local models will handle most daily tasks within five years.

  3. Mastra BlogOfficialAI score67

    Mastra launches Agent Controller GA, a runtime for long-running agent sessions

    AIMastra has released Agent Controller in general availability, a runtime that hosts long-running agent sessions around the agent loop. The team says it was first built for Mastra Code and expanded to support Mastra Factory, which runs many concurrent sessions, and that memory usage in long-running Mastra Code processes dropped from 2–20 GB to 300–750 MB after optimizing UI state snapshots.

    Why it matters: The post explains how the controller evolved from one developer's session to many concurrent sessions, with measured memory and storage changes useful to engineers building multi-user agent apps.

  4. Claude BlogOfficialAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    AIComcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    Why it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.

Oct 5

Oct 5Mon
  1. ThariqXAI score22

    Thariq says HTML planning is more token efficient than raw HTML

    AIThariq says planning with HTML is much more token efficient than generating raw HTML. The model does not need to recreate components or logic for common elements such as state machines, diagrams, and code snippets. Background from the quoted post says he is building a Claude Code skill that generates HTML plans, with linting to reduce common failures.

  2. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  3. Dongxi NLPXAI score60

    Reflection AI's Beam open model is compared against leading Chinese models

    AIThe author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.

  4. NVIDIA AIOfficialAI score39

    NVIDIA releases Nemotron-Labs-3-Competitive-Coding model on Hugging Face

    AINVIDIA has published Nemotron-Labs-3-Competitive-Coding on Hugging Face, a competitive-programming specialist model built on Nemotron-3-Ultra. The model is available in the NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 repository, indicating a 550B-parameter total size with 55B active parameters in NVFP4 format.

  5. ThariqXAI score22

    Thariq shares a Claude Code skill for generating better HTML plans

    AIThariq, who works at Anthropic, is developing a skill for Claude Code that produces HTML plans using simple language, code snippets, surfaced questions, and mockups. Linting is used to reduce common failure cases Claude encounters, and he is seeking feedback before a broader release.

    Video from @trq212's post
  6. ReflectionOfficialAI score23

    Reflection AI's Beam model pretrained in four weeks on 24T tokens

    AIReflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities. The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class. It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

    Image from @reflection_ai's post
  7. CursorOfficialAI score39

    Cursor lets users replace its system prompt with their own

    AICursor is enabling an option to replace its built-in system prompt with a custom one, while rules, skills, and tool schemas still load. The feature is being rolled out account by account rather than to all users at once.

    Image from @cursor_ai's post
  8. Together AIOfficialAI score34

    Together AI launches Together Link to run open models in coding harnesses

    AITogether AI has announced Together Link, which lets developers run frontier open models inside their favorite coding harness. The product includes spending tracking and an Auto router that selects low-cost models for quick fixes and more capable models for harder tasks.

    Video from @togethercompute's post
  9. GitHub Blog · AI & MLOfficialAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  10. FireworksOfficialAI score34

    DeepSeek V4.1 Flash now available for training on Fireworks

    AIFireworks AI has made DeepSeek V4.1 Flash available for training on its Dedicated Training API and Managed Training surfaces. The post positions the model as a strong base for agentic coding, terminal automation, and tool use, and notes it is cost-efficient to serve.

  11. Replit ⠕OfficialAI score22

    Replit weekly changelog adds GPT-6.1 Sol and Claude Sonnet 5.5 models

    AIReplit shipped a weekly update letting users build with GPT-6.1 Sol and Claude Sonnet 5.5, along with an Ask agent integration with Jev. The release also includes an updated Settings UI and enterprise Workplace controls for company-wide rules and controlled exceptions. Full details are in the Replit changelog.

  12. clem 🤗XAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

    Why it matters: The capture proxy lets models train inside real coding harnesses without reimplementing them, with measured gains and a comparison against SFT on the same data.

    Image from @ClementDelangue's post
  13. Guillermo RauchXAI score44

    gdp-ts brings compile-time authorization proofs to TypeScript APIs

    AIGuillermo Rauch introduced gdp-ts, a library, linter, and AI skill that uses "proofs" to enforce that sensitive functions are called only after an authorization check. The TypeScript typechecker verifies these proofs at compile time, aiming to stop security bugs from shipping, including those written by AI agents. The README models a Vercel API constraint requiring a role and entitlement proof to change a Project's password.

    Video from @rauchg's post
  14. Cloudflare Blog · AIOfficialAI score40

    Cloudflare Birthday Week 2026 unveils cf CLI, EmDash CMS, and post-quantum tools

    AICloudflare announced 46 products and updates during Birthday Week 2026, including the cf CLI for the entire Cloudflare API and EmDash, an open-source Astro-based serverless CMS whose plugins run in isolated Worker sandboxes. The company also said it plans to become a public certificate authority that issues free Merkle Tree Certificates for post-quantum authentication.

  15. meng shaoXAI score47

    Emil Kowalski's /break-ui Skill Stress-Tests UIs With Realistic Worst-Case Data

    AIThe /break-ui Skill, added to the Skills For Designers and Engineers repo with 43K stars and 1.9M installs, plays the most annoying real user to stress UI components with worst-case but realistic data. It targets bugs manual testing misses, such as "1 members" pluralization errors, zero-value "0 seconds ago" rendering, cross-timezone date shifts, and emoji or CJK names breaking initials logic. The skill reports issues before fixing them, and only changes the data, never the component.

    Image from @shao__meng's post

Oct 4

Oct 4Sun
  1. Together AI BlogOfficialAI score38

    Together Link Routes Coding Agents to Open Models, Cutting Spend Over 50%

    AITogether Link connects coding agents such as Claude Code, Codex, OpenCode, and Pi to open models on Together AI, which the company says cuts spend by over 50%. Setup takes one command, and its "Auto" mode routes each session's first task to a fast low-cost model or a frontier model, with a per-session tracker comparing costs against Opus 5.5.

  2. Epoch AIOfficialAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

  3. Guillermo RauchXAI score46

    Vercel's Guillermo Rauch says Turborepo moved from Go to Rust

    AIVercel completed migrating Turborepo from Go to Rust, which Rauch says was chosen for better low-level OS access despite controversial returns on human migration costs. He argues that what is best for humans is no longer necessarily best for business now that agents are writing code, and suggests Rust may not be the final toolchain.

  4. Jerry LiuXAI score23

    Jerry Liu Says ChatGPT/Codex Offers Best Agent Interface for Deep Work

    AIJerry Liu says ChatGPT/Codex currently has the best agent interface for deep work, unifying coding and knowledge work in one place with forking support that Claude's app lacks. He still prefers Claude Code CLI as close to the best a CLI can be, and uses Opus 5.5 mainly through it for product demos, while noting a GUI is sometimes nicer.

  5. KhazixXAI score45

    Claude Opus 5.5 weekly quota outlasts GPT-6 Astra by tenfold

    AIThe author tracked token usage over three days and estimated that a $200 Claude plan delivers about $3,400 of API-equivalent value per week, versus about $1,700 for a $200 Codex plan. With cache hit rates of 98.94% for Claude Code and 98.34% for Codex, the author says GPT-6 Astra costs roughly five times more than Claude Opus 5.5, making the Claude weekly quota last about ten times longer.

    Image from @Khazix0918's post
  6. Harrison ChaseXAI score31

    LangChain cuts coding agent costs with tracking, caps, and routing

    AILangChain says its coding agent costs fell significantly for a second straight month after adopting three steps. The steps are cost visibility through LangSmith tracing, per-user cost caps via its LLM gateway, and harness optimization such as model routing in its open-source OpenSWE cloud agent harness.

    Image from @hwchase17's post
  7. Yuchen JinXAI score23

    Yuchen Jin says AI agents are replacing terminals as the coding interface

    AIYuchen Jin argues that terminals, built around files, commands, and processes, are giving way to AI agents where users state intent and the agent operates the machine. He says understanding Linux and systems fundamentals remains valuable as a moat. In a follow-up, he calls the terminal era over for coding agents, saying persistent context matters more than tabs, and names the Codex desktop app as the best agentic UI for now.

Oct 3

Oct 3Sat
  1. ClineOfficialAI score35

    Ling 3.1 Flash is available free in Cline until October 13

    AICline says Ling 3.1 Flash is now available in its platform and free through October 13. The 560B total parameter mixture-of-experts model activates 25B parameters and is described as on par with open-weights models Kimi K3 and DeepSeek V4 Pro.

    Image from @cline's post