Skip to contentSkip to stories

Updated

Coding

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. Mastra BlogOfficialAI score67

    Mastra launches Agent Controller GA, a runtime for long-running agent sessions

    AIMastra has released Agent Controller in general availability, a runtime that hosts long-running agent sessions around the agent loop. The team says it was first built for Mastra Code and expanded to support Mastra Factory, which runs many concurrent sessions, and that memory usage in long-running Mastra Code processes dropped from 2–20 GB to 300–750 MB after optimizing UI state snapshots.

    Why it matters: The post explains how the controller evolved from one developer's session to many concurrent sessions, with measured memory and storage changes useful to engineers building multi-user agent apps.

  2. Claude BlogOfficialAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    AIComcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    Why it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.

Oct 5

Oct 5Mon
  1. ThariqXAI score22

    Thariq says HTML planning is more token efficient than raw HTML

    AIThariq says planning with HTML is much more token efficient than generating raw HTML. The model does not need to recreate components or logic for common elements such as state machines, diagrams, and code snippets. Background from the quoted post says he is building a Claude Code skill that generates HTML plans, with linting to reduce common failures.

  2. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  3. Dongxi NLPXAI score60

    Reflection AI's Beam open model is compared against leading Chinese models

    AIThe author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.

  4. NVIDIA AIOfficialAI score39

    NVIDIA releases Nemotron-Labs-3-Competitive-Coding model on Hugging Face

    AINVIDIA has published Nemotron-Labs-3-Competitive-Coding on Hugging Face, a competitive-programming specialist model built on Nemotron-3-Ultra. The model is available in the NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 repository, indicating a 550B-parameter total size with 55B active parameters in NVFP4 format.

  5. ClineOfficialAI score19

    Cline launches a desktop app alongside its CLI

    AICline says its CLI can be installed globally with npm i -g cline, and it has also released a new Desktop app. The company notes several free model promotions are available to try the Desktop app.

  6. ThariqXAI score22

    Thariq shares a Claude Code skill for generating better HTML plans

    AIThariq, who works at Anthropic, is developing a skill for Claude Code that produces HTML plans using simple language, code snippets, surfaced questions, and mockups. Linting is used to reduce common failure cases Claude encounters, and he is seeking feedback before a broader release.

    Video from @trq212's post
  7. ReflectionOfficialAI score23

    Reflection AI's Beam model pretrained in four weeks on 24T tokens

    AIReflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities. The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class. It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

    Image from @reflection_ai's post
  8. dexXAI score16

    Dex Horthy argues human review still gives AI work its edge

    AIDex Horthy argues there will always be an advantage in reviewing AI output, since unreviewed work tends toward generic "AI slop." He adds that the reviewed object may not always be code, and speculates that once models far surpass humans, human interference could make results worse.

  9. Google AntigravityOfficialAI score16

    Artist Leo Villareal used Gemini and Google Antigravity for light sculpture software

    AIArtist Leo Villareal used Gemini and Google Antigravity to build the authoring software behind his light sculpture "Buckyball," which debuted in New York City this summer. Google Cloud's post frames the project as a case study of AI tools supporting creative work. The source provides no further technical details about the software or results.

  10. CursorOfficialAI score39

    Cursor lets users replace its system prompt with their own

    AICursor is enabling an option to replace its built-in system prompt with a custom one, while rules, skills, and tool schemas still load. The feature is being rolled out account by account rather than to all users at once.

    Image from @cursor_ai's post
  11. Together AIOfficialAI score34

    Together AI launches Together Link to run open models in coding harnesses

    AITogether AI has announced Together Link, which lets developers run frontier open models inside their favorite coding harness. The product includes spending tracking and an Auto router that selects low-cost models for quick fixes and more capable models for harder tasks.

    Video from @togethercompute's post
  12. GitHub Blog · AI & MLOfficialAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  13. FireworksOfficialAI score34

    DeepSeek V4.1 Flash now available for training on Fireworks

    AIFireworks AI has made DeepSeek V4.1 Flash available for training on its Dedicated Training API and Managed Training surfaces. The post positions the model as a strong base for agentic coding, terminal automation, and tool use, and notes it is cost-efficient to serve.

  14. Replit ⠕OfficialAI score22

    Replit weekly changelog adds GPT-6.1 Sol and Claude Sonnet 5.5 models

    AIReplit shipped a weekly update letting users build with GPT-6.1 Sol and Claude Sonnet 5.5, along with an Ask agent integration with Jev. The release also includes an updated Settings UI and enterprise Workplace controls for company-wide rules and controlled exceptions. Full details are in the Replit changelog.

  15. clem 🤗XAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

    Why it matters: The capture proxy lets models train inside real coding harnesses without reimplementing them, with measured gains and a comparison against SFT on the same data.

    Image from @ClementDelangue's post
  16. Guillermo RauchXAI score44

    gdp-ts brings compile-time authorization proofs to TypeScript APIs

    AIGuillermo Rauch introduced gdp-ts, a library, linter, and AI skill that uses "proofs" to enforce that sensitive functions are called only after an authorization check. The TypeScript typechecker verifies these proofs at compile time, aiming to stop security bugs from shipping, including those written by AI agents. The README models a Vercel API constraint requiring a role and entitlement proof to change a Project's password.

    Video from @rauchg's post
  17. Cloudflare Blog · AIOfficialAI score40

    Cloudflare Birthday Week 2026 unveils cf CLI, EmDash CMS, and post-quantum tools

    AICloudflare announced 46 products and updates during Birthday Week 2026, including the cf CLI for the entire Cloudflare API and EmDash, an open-source Astro-based serverless CMS whose plugins run in isolated Worker sandboxes. The company also said it plans to become a public certificate authority that issues free Merkle Tree Certificates for post-quantum authentication.

  18. meng shaoXAI score47

    Emil Kowalski's /break-ui Skill Stress-Tests UIs With Realistic Worst-Case Data

    AIThe /break-ui Skill, added to the Skills For Designers and Engineers repo with 43K stars and 1.9M installs, plays the most annoying real user to stress UI components with worst-case but realistic data. It targets bugs manual testing misses, such as "1 members" pluralization errors, zero-value "0 seconds ago" rendering, cross-timezone date shifts, and emoji or CJK names breaking initials logic. The skill reports issues before fixing them, and only changes the data, never the component.

    Image from @shao__meng's post

Oct 4

Oct 4Sun
  1. Together AI BlogOfficialAI score38

    Together Link Routes Coding Agents to Open Models, Cutting Spend Over 50%

    AITogether Link connects coding agents such as Claude Code, Codex, OpenCode, and Pi to open models on Together AI, which the company says cuts spend by over 50%. Setup takes one command, and its "Auto" mode routes each session's first task to a fast low-cost model or a frontier model, with a per-session tracker comparing costs against Opus 5.5.

  2. Epoch AIOfficialAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

  3. Boris PowerXAI score13

    OpenAI signals a rapid run of Codex and work-user improvements

    AIBoris Power, who is linked to OpenAI, posted "Time to 🚢 🚢 🚢!" to signal an imminent round of shipments. Quoted background from @thsottiaux says that over the next 28 days the team will ship one clear improvement relevant to most Codex and work users each day, or a full reset.

  4. Guillermo RauchXAI score46

    Vercel's Guillermo Rauch says Turborepo moved from Go to Rust

    AIVercel completed migrating Turborepo from Go to Rust, which Rauch says was chosen for better low-level OS access despite controversial returns on human migration costs. He argues that what is best for humans is no longer necessarily best for business now that agents are writing code, and suggests Rust may not be the final toolchain.

  5. Jerry LiuXAI score23

    Jerry Liu Says ChatGPT/Codex Offers Best Agent Interface for Deep Work

    AIJerry Liu says ChatGPT/Codex currently has the best agent interface for deep work, unifying coding and knowledge work in one place with forking support that Claude's app lacks. He still prefers Claude Code CLI as close to the best a CLI can be, and uses Opus 5.5 mainly through it for product demos, while noting a GUI is sometimes nicer.