Skip to contentSkip to stories

Updated

#Coding

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Latent.SpaceAI score55

    Anthropic's Thariq explains what Claude Code mods can access and control

    AIAnthropic's Thariq describes Claude Code mods, which can read conversation scope such as turn count and token usage. Mods run in process, so they can spawn subagents, parse their results with structured output, and modify the UI, which hooks cannot do. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app.

    Video from @latentspacepod's post
  2. Lydia Hallie ✨AI score62

    Claude Code adds mods that customize behavior and UI via TypeScript plugins

    AIClaude Code can now be modified with mods that change its behavior, customize the UI, and add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app. A TypeScript function can intercept internal events such as tool calls, prompts, model requests, and renders, and add custom UI and commands.

  3. Comfy BlogAI score47

    Hakoniwa uses Comfy Agent to make the animated short YUI

    AIArtist 852 Hakoniwa made YUI, described as the first animated short created with Comfy Agent, which the ComfyUI team says took three days of focused work by one person at about 200,000 yen in total cost, excluding labor. The source says the film was made mostly with Seedance 2.5 and the making-of video with MiniMax H3, with Comfy Agent used to regenerate shots and compare video models.

  4. Lewis Tunstall @ COLM 🌉AI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  5. JetBrains AI BlogAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  6. Anthropic NewsroomAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  7. Manus BlogAI score47

    Manus 2.0 Adds Game Dev for Building Multiplayer Games Without Coding

    AIManus has launched Game Dev in Manus 2.0, a feature that lets users with no coding experience build games with a real-time tweak panel, asset management, and multiplayer servers. The tweak panel lets users adjust settings such as speed, gravity, damage, and spawn rate while playing. Manus also handles much of the multiplayer infrastructure, including server deployment and networking, so games can be shared and played with friends.

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  2. Apple Machine Learning ResearchAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  3. Varun MohanAI score40

    Google announces Gemini 4 Argon, a new frontier model for software tasks

    AIGoogle announced Gemini 4 Argon, a new frontier model that delivers frontier performance across complex software tasks, according to Varun Mohan. Thousands of Googlers have been using it internally in Antigravity, and it is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

  4. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  5. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  6. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.