Skip to contentSkip to stories

Updated

Coding

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. PyTorch BlogOfficialAI score38

    Meta's Jagged Flash Attention kernel beats FA4 on Blackwell using TLX

    AIPyTorch Blog says its Jagged Flash Attention kernel, the attention kernel behind Meta's Generative Ads Model, runs on NVIDIA B200 Blackwell built with TLX (Triton Low-level Extensions). The kernel is about 3.2K lines, roughly 3× less code than the roughly 10K-line CuteDSL FlashAttention-4 (FA4) kernel, and outperforms FA4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass.

  2. ClineOfficialAI score34

    Cline reports DeepSeek V4 Pro costs about 30x less than Claude Opus 5

    AICline says two of its largest tasks in the last 30 days each processed 9B tokens, costing about $8,500 on Claude Opus 5 versus about $300 on DeepSeek V4 Pro. The post says that is roughly 30x cheaper for the same token count, letting users run large-horizon work without spending thousands or waiting on limit resets.

  3. Boris ChernyXAI score54

    Claude Code adds mods that customize behavior and UI via plugins

    AIClaude Code now supports mods that change its behavior, customize the UI, or add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app, and the author says mods can be shared so others can try them.

  4. Latent.SpaceXAI score55

    Anthropic's Thariq explains what Claude Code mods can access and control

    AIAnthropic's Thariq describes Claude Code mods, which can read conversation scope such as turn count and token usage. Mods run in process, so they can spawn subagents, parse their results with structured output, and modify the UI, which hooks cannot do. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app.

    Video from @latentspacepod's post
  5. Lydia Hallie ✨XAI score62

    Claude Code adds mods that customize behavior and UI via TypeScript plugins

    AIClaude Code can now be modified with mods that change its behavior, customize the UI, and add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app. A TypeScript function can intercept internal events such as tool calls, prompts, model requests, and renders, and add custom UI and commands.

    Why it matters: The post details how mods hook into Claude Code's tool calls, prompts, and rendering, which shows what extending the coding agent actually involves.

  6. Comfy BlogOfficialAI score47

    Hakoniwa uses Comfy Agent to make the animated short YUI

    AIArtist 852 Hakoniwa made YUI, described as the first animated short created with Comfy Agent, which the ComfyUI team says took three days of focused work by one person at about 200,000 yen in total cost, excluding labor. The source says the film was made mostly with Seedance 2.5 and the making-of video with MiniMax H3, with Comfy Agent used to regenerate shots and compare video models.

  7. DatabricksOfficialAI score20

    Databricks Smart Routing assigns each coding task to a suitable model

    AIDatabricks' Smart Routing evaluates each coding task separately and selects the lowest-cost model capable of handling it, balancing quality, latency, and cost. In a demo, Omnigent splits an app build into planning, backend, and frontend work, routes each part to a different model, and runs some tasks in parallel.

    Video from @databricks's post
  8. Josh WoodwardXAI score34

    Google launches Stitch CLI to generate design ideas from terminal

    AIGoogle has introduced the @google/stitch CLI, letting users generate screens and design systems without leaving the terminal. It connects to local coding agents and can send a local dev server snapshot to Stitch. The tool complements the existing Stitch MCP and SDK, and can also be driven through agents such as Antigravity.

  9. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  10. OpenRouter · New modelsBlogAI score36

    Pareto 26.10 Preview: A Multimodal Model for Research, Coding and Agents

    AIPareto 26.10 Preview is a multimodal composite model built for research, coding, and agentic workflows. It is described as delivering frontier-level performance across a broad range of general-purpose tasks, though the source excerpt is a preview and provides no benchmark scores, parameter counts, pricing, or availability details.

  11. JetBrains AI BlogOfficialAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  12. Anthropic NewsroomOfficialAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  13. Manus BlogOfficialAI score47

    Manus 2.0 Adds Game Dev for Building Multiplayer Games Without Coding

    AIManus has launched Game Dev in Manus 2.0, a feature that lets users with no coding experience build games with a real-time tweak panel, asset management, and multiplayer servers. The tweak panel lets users adjust settings such as speed, gravity, damage, and spawn rate while playing. Manus also handles much of the multiplayer infrastructure, including server deployment and networking, so games can be shared and played with friends.

Sep 30

Sep 30Wed
  1. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  2. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  3. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  4. koray kavukcuogluXAI score62

    Google's Koray Kavukcuoglu Announces Gemini 4 Argon for Trusted Defenders First

    AIGoogle is sharing Gemini 4 Argon first with trusted defenders in its Fairwind Program, with frontier capabilities in coding, knowledge work, and cyber security defense. The model is also being rolled out to the US government, and the company plans broader availability as testing progresses and safeguards allow.

    Why it matters: The author is a Google DeepMind leader announcing the model directly, so the rollout limits to trusted defenders and government are the key detail to note.

  5. Varun MohanXAI score40

    Google announces Gemini 4 Argon, a new frontier model for software tasks

    AIGoogle announced Gemini 4 Argon, a new frontier model that delivers frontier performance across complex software tasks, according to Varun Mohan. Thousands of Googlers have been using it internally in Antigravity, and it is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

  6. Google AIOfficialAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  7. Google DeepMindOfficialAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  8. Microsoft CopilotOfficialAI score23

    Microsoft's new Copilot combines Home, Code, and Autopilot in one app

    AIMicrosoft's new Copilot brings Home, Code, and Autopilot together in one place for creating custom apps, building decks, automating workflows, and resuming work. Users can start using the Copilot app now and try new features as they become available in Frontier.

    Image from @MSFTCopilot's post
  9. Google · Gemini appOfficialAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  10. eric zakariassonXAI score22

    Cursor promotes engineering bots in Grok bot marketplace

    AICursor's Eric Zakariasson announced engineering bots available in the Grok bot marketplace on x.ai. The bots can hand off coding tasks to Cursor, manage pull requests through GitHub and Origin plugins, and share video demos of what they build.

    Image from @ericzakariasson's post
  11. FireworksOfficialAI score34

    GLM 5.3 Flash now available for training on Fireworks' Serverless API

    AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.

  12. ClineOfficialAI score34

    Cline desktop app can run agents on remote Linux servers over SSH

    AICline says its desktop app can keep running on a laptop while the agent executes on any Linux machine reachable via SSH, configured under Settings → Remote. The setup requires no root access, no npm, and no public port, uploading a self-contained helper and tunneling only the authenticated Cline protocol.

  13. GitHubOfficialAI score42

    Project HydraFusion Now Available in GitHub Copilot App and VS Code

    AIProject HydraFusion is now available in the GitHub Copilot app and VS Code, where users can select it like any other model. Behind the scenes, it routes a task across multiple models to draft, critique, revise, or escalate, then returns one combined result.

    Video from @github's post
  14. Ant LingOfficialAI score31

    Ant Ling model turns plain-language prompts into interactive Three.js pages

    AIAnt Ling can convert plain-language prompts into standalone, interactive Three.js pages covering topics such as an internal combustion engine, an optical-disc reader, paramecium organelles, and vector-field divergence. The output is runnable code rather than just an explanation.

    Video from @AntLingAGI's post