Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 13

Aug 13Thu
  1. Meituan LongCatOfficialAI score46

    LongCat-2.0 Free for One Week on Nous Portal with Hermes Agent

    AILongCat-2.0, Meituan's model, is now live on the Nous Portal and free to try with Hermes Agent for one week. Nous describes it as a 1.6T-parameter MoE with a 1M context built for agentic coding, scoring 70.8 on Terminal-Bench 2.1. It can ingest an entire codebase in one pass, and the Portal is at

  2. Matei ZahariaXAI score44

    Databricks adds Smart Routing to Unity AI Gateway for coding agents

    AIDatabricks has made Smart Routing available in Unity AI Gateway to improve coding agent quality and cost. It matches each coding task to a suitable model and harness based on task needs while preserving good cache hit rates. Databricks says this can match frontier quality while cutting task costs by 30% or more.

  3. Google AI DevelopersOfficialAI score47

    Gemini 3.7 Flash generates native apps across five mobile frameworks

    AIGoogle's Gemini 3.7 Flash in Antigravity turns one architecture spec into production-ready code for five mobile frameworks: Flutter, SwiftUI, Jetpack Compose, React Native, and NativeScript. The post presents this as generating native apps without boilerplate from a single specification.

    Video from @googleaidevs's post
  4. Varun MohanXAI score52

    Gemini 3.7 Flash Goes Live in Google Antigravity for Coding and Agents

    AIGoogle has made Gemini 3.7 Flash available in Google Antigravity, described as its most intelligent workhorse model yet for coding and agents. Users can download or upgrade Antigravity to try the model. The author, Varun Mohan, says the model brings a big capability improvement at half the API cost.

  5. Augment Code BlogOfficialAI score44

    Augment Code Expands AI Review Loop to Automate PR-to-Merge Workflow

    AIAugment Code describes an expanded AI-native review system in which specialized agents handle review, repair, and verification from pull request to merge. Humans still make judgment calls and the final merge decision, with the company claiming a 3× increase in code output in its earlier review system.

  6. Google AI DevelopersOfficialAI score75

    Google releases Gemini 3.7 Flash for coding and agentic tasks

    AIGoogle AI Developers announced Gemini 3.7 Flash as its most intelligent workhorse model yet for coding and agents, citing higher instruction adherence, first-pass code accuracy, and high-quality agentic execution. The post shows the model building a complex 3D web game in Antigravity, covering Three.js engine logic, asset orchestration with PBR textures and Nano Banana sprite sheets, and procedural sound effects.

    Why it matters: The post shows a concrete build workflow across engine logic, assets, and audio, which helps readers judge how the model handles multi-step agentic coding.

    Video from @googleaidevs's post
  7. Demis HassabisXAI score67

    Google releases Gemini 3.7 Flash with coding and web development upgrades

    AIGoogle DeepMind has released Gemini 3.7 Flash, which the post says is stronger for coding, knowledge work, and web development. Its introductory price is half the original cost of Gemini 3.6 Flash.

    Why it matters: The post names concrete upgrade areas and a price change against the prior version, which helps readers compare it with earlier Flash releases.

  8. koray kavukcuogluXAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

    Image from @koraykv's post
  9. DeepSeek API NewsOfficialAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. Factory NewsOfficialAI score40

    Factory Launches Agent Effectiveness to Link Droid Usage to Delivery Outcomes

    AIFactory's Agent Effectiveness, now in Private Preview within Factory Analytics, connects Droid sessions to cycle time, work intent, and shipped artifacts drawn from project, issue-tracking, and source control tools. Its Throughput, Output, and Attribution views show where delivery is speeding up, how spend splits across feature, maintenance, bug-fixing, and exploration work, and which projects and issues the output maps to. Admins enable it by connecting Jira, Linear, GitHub, or GitLab and turning on the Advanced Analytics enterprise control.

  2. Cursor ChangelogOfficialAI score42

    Cursor Cloud Agents Start 3x Faster With Builds

    AICursor's Cloud Agents now start from prebuilt copies of development environments, cutting startup time by 3x, with environments booting 10x faster internally and 3x faster time to first token. Builds are included at no additional cost, and failed builds are not activated, so agents keep using the last successful build while users debug in the background.

Aug 11

Aug 11Tue
  1. Zed BlogOfficialAI score72

    Zed introduces Delta, a multiplayer environment for coding with agents

    AIand reviewing their code, and invites first users into a private beta. Delta keeps code and conversations connected through DeltaDB, which captures edits and conversations between git commits and works with existing repositories. The app also supports cloud runners, browser-based sharing, and live syncing of Claude Code sessions.

    Why it matters: The post explains how the new Delta app links conversations with code history, which clarifies a shift in how teams review agent-written changes.

  2. Z.aiOfficialAI score33

    ZCode reaches 1 million users and resets GLM Coding Plan limits

    AIZ.ai says its ZCode platform has reached 1 million users, and it has reset usage limits for all GLM Coding Plan users as a thank-you. The post also announces an update aimed at turning long-horizon capabilities into completed engineering work, reporting a 98% cache hit rate that provides around 1.8x more usage.

    Video from @Zai_org's post

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceOfficialAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. Matei ZahariaXAI score44

    Matei Zaharia says AI Gateways let teams cut token costs centrally

    AIMatei Zaharia argues AI tokens are now a resource to optimize in software engineering, with companies routing all AI usage through an AI Gateway. The approach enables centralized analysis, which found settings on Claude Code and Codex that can substantially lower cost, plus smart routing and per-task budgets for engineers.

Aug 5

Aug 5Wed
  1. Meituan LongCatOfficialAI score34

    LongCat-2.0 is now free on OpenCode for coding

    AIMeituan's LongCat-2.0, a coding-optimized model with 1M context and full open-source availability, is now free to use on OpenCode. The announcement comes from Meituan LongCat's official X account and does not include pricing details beyond the free offer.

  2. Prime Intellect BlogOfficialAI score75

    Prime Agent launches open-source self-improving RLM coding harness

    AIPrime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.

    Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.

Aug 4

Aug 4Tue
  1. Zed BlogOfficialAI score65

    Zed Enables OS-Level Sandboxing by Default for Its Agent Panel

    AIZed's agent panel now sandboxes its terminal and fetch tools by default, starting in release 1.14, and the restrictions are enforced by the operating system rather than by agent instructions. By default the sandbox blocks writes outside project directories, writes to .git, and network requests, and agents can request temporary escalation with a stated reason. The post also notes that sandboxing covers only those tools and does not protect against other tools, external programs, or the regular built-in terminal.

    Why it matters: The post explains how OS-enforced sandboxing limits agent terminal and fetch access, and why fine-grained command rules fall short of it.

  2. Mckay WrigleyXAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

Aug 3

Aug 3Mon
  1. JetBrains AI BlogOfficialAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

Aug 2

Aug 2Sun
  1. OpenRouter BlogOfficialAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Aug 1

Aug 1Sat
  1. Andrej KarpathyXAI score66

    Karpathy tests Opus 5 by rendering Lord of the Rings opening in 3D

    AIAndrej Karpathy gave Claude Opus 5 the first paragraph of Lord of the Rings with a 1M token budget and asked for a Three.js render. Opus spent about two hours writing 5500 lines of code that procedurally renders the story, which Karpathy calls janky but fun. He notes the model struggled to audit its work because it cannot efficiently perceive video or play the resulting game, relying on slow screenshots that led to several errors.

    Video from @karpathy's post

Jul 31

Jul 31Fri
  1. DeepSeekOfficialAI score42

    DeepSeek-V4-Flash API launches in public beta with stronger agent performance

    AIDeepSeek has released the official DeepSeek-V4-Flash API in public beta, with substantially upgraded agent capabilities. The company says its benchmark scores now far surpass those of V4-Pro-Preview. The official V4-Flash natively supports the Responses API format and is adapted for Codex, with configuration details in DeepSeek's API docs.

    Image from @deepseek_ai's post
  2. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. Thinking MachinesOfficialAI score44

    Thinking Machines' Inkling-Small Gains From Lessons of Larger Inkling

    AIThinking Machines says Inkling-Small began training after its larger counterpart and benefits from the lessons learned. The post cites an improved pre-training data mix, a refined ML recipe, on-policy distillation using Inkling as the teacher, and two additional weeks of agentic coding RL.

Jul 29

Jul 29Wed
  1. Air Street PressBlogAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

Jul 28

Jul 28Tue
  1. Augment Code BlogOfficialAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. JetBrains AI BlogOfficialAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 27

Jul 27Mon
  1. Kimi.aiOfficialAI score52

    Kimi K3 launches on Together AI as a Day 0 partner

    AIKimi K3 is now available on Together AI, which is a Day 0 launch partner for the model. Together AI offers developers immediate access to K3 through high-throughput inference aimed at coding agents and production workloads.

    Image from @Kimi_Moonshot's post

Jul 26

Jul 26Sun
  1. Fireworks AI BlogOfficialAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

  2. Philipp SchmidBlogAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

Jul 25

Jul 25Sat

Jul 24

Jul 24Fri
  1. Noah ZwebenXAI score44

    Noah Zweben shares a favorite Opus 5 anecdote from his TA days

    AIAnthropic's Noah Zweben says a tornado-physics assignment he once TA'd for, built in Unity, is his favorite Opus 5 example so far. The quoted Atomic Chat post compares Opus 5, Fable 5, Kimi K3, and GPT 5.6 on three HTML physics scenes, with Opus 5 costing $1.40 versus Fable 5's $2.82.