Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 2

Sep 2Wed
  1. Varun MohanXAI score57

    Gemini 3.8 Flash released with gains in agentic coding and knowledge work

    AIGoogle's Gemini 3.8 Flash is out, and Varun Mohan says it substantially improves on 3.7 Flash for agentic coding and general knowledge work. It is now available to everyone on Antigravity. The attached benchmark table lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with introductory pricing through December 31, 2026.

    Image from @_mohansolo's post
  2. Google AI DevelopersOfficialAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post
  3. Logan KilpatrickXAI score42

    Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark

    AIGoogle's Gemini 3.8 Flash reached 73.7% on the DeepSWE 1.1 benchmark, according to a post from Logan Kilpatrick. The post gives no further details on methodology, comparisons, or pricing.

    Image from @OfficialLoganK's post
  4. koray kavukcuogluXAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Why it matters: The chart compares Gemini 3.8 Flash's DeepSWE score and average cost per task against several frontier models, showing where the cost-performance tradeoff lands.

    Image from @koraykv's post
  5. Logan KilpatrickXAI score62

    Google releases Gemini 3.8 Flash with gains in agentic and coding tasks

    AIGoogle announced Gemini 3.8 Flash, its third updated Flash model in six weeks, citing improvements in agentic and coding capabilities. The benchmark table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing of $1.50 and $7.50 expiring December 31, 2026. Terminal-bench 2.1 shows 89.4% for Gemini 3.8 Flash against 85.8% for Gemini 3.7 Flash.

    Why it matters: The benchmark table compares Gemini 3.8 Flash against Gemini 3.7 Flash and rival models, showing where the gains and remaining gaps fall across coding and agent tasks.

    Image from @OfficialLoganK's post
  6. Google AI StudioOfficialAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

    Why it matters: The post gives specific pricing and access paths for the new model, letting developers compare it with 3.7 Flash on cost and availability.

    Image from @GoogleAIStudio's post

Sep 1

Sep 1Tue
  1. Meituan LongCatOfficialAI score46

    LongCat-2.0 Now Free to Try in Cline

    AIMeituan's LongCat-2.0, a 1.6T open-weights MoE model with a 1M context window, is now free to use in Cline. Cline's post says it scores similarly to Claude Opus 4.7 and Gemini 3.1 Pro. Users can select it under free models via /model after installing Cline with npm i -g cline.

  2. catXAI score50

    Anthropic's Claude Fable 5.1 enables more ambitious, months-long projects

    AIAnthropic's team says Claude Fable 5.1 has let them take on projects that previously would have taken months, and invites users to try it in Claude Code, Claude Cowork, and Claude Tag. The post asks what big bets users want to make, and it builds on Anthropic's announcement of Claude Fable 5.1 and Claude Mythos 5.1 as its most advanced models for coding and knowledge work.

  3. Anthropic · YouTubeOfficialAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesOfficialAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. Philipp SchmidBlogAI score60

    Frontier models now compose Bash workflows that replace dedicated coding tools

    AIThe author rebuilt an agent harness with only a bash tool and a media viewer, and task completion stayed in the same range. Three example workflows show multi-file edits, bisecting a flaky test, and correlating compressed logs in SQLite, with the intermediate data kept out of the model context. In a comparison against separate file, edit, and search tools on the same coding tasks, the shell-centered setup performed on par or better, though the author notes images still need a multimodal channel.

Aug 29

Aug 29Sat
  1. Tencent HyOfficialAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri
  1. Thomas DohmkeXAI score25

    Entire launches one API for code and coding sessions

    AIEntire positions itself as a unified API for code and coding sessions, working across any agent, repo, and session as a coding system of record. The post frames this as a single interface layer for coding work, comparing it to unified-interface products in payments, models, and banking.

  2. TinkerOfficialAI score52

    GLM-5.3 from Z.ai is now available on Tinker with 256k context

    AITinker announces that Z.ai's GLM-5.3 is now available on its platform with a 256k context window. Tinker says it is currently the strongest open-weights model on coding evals including Terminal-Bench 3.0 and DeepSWE 1.1, built on the same base as GLM-5.2 with scaled-up post-training.

  3. Andrew NgXAI score20

    Andrew Ng maps software engineering fundamentals for agentic coding era

    AIAndrew Ng published an AI Engineering Skills map covering the software engineering fundamentals developers need when working with coding agents. He argues that understanding full-stack architecture, data management, system design, security, reliability, and production scaling lets developers steer agents toward the right tradeoffs in latency, availability, consistency, and cost. Without these fundamentals, vibe-coded applications often end up with poor tradeoffs the developer never anticipated.

  4. Z.aiOfficialAI score62

    Z.ai releases GLM-5.3 as open-weight model for agentic coding and cyber defense

    AIZ.ai has made GLM-5.3 open-weight, so users can download, run, and customize its weights. The company describes it as its most capable model for agentic coding and cyber defense, with weights on Hugging Face and details in a tech blog.

    Why it matters: The source ties the open-weight release to coding, agentic, and cyber defense use, with a weights link for anyone who wants to run or customize it.

    Image from @Zai_org's post

Aug 27

Aug 27Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

Aug 26

Aug 26Wed
  1. Cursor ChangelogOfficialAI score46

    Cursor Cloud Agents now let you start projects from scratch without a repo

    AICursor Cloud Agents no longer require a connected GitHub or other third-party SCM provider to begin work. Users select "Start from scratch" in the repo picker, and Cursor creates an Origin repo in the background that can be saved as a private or internal repo via "Create repo." Cursor also now port-forwards the cloud agent's live environment to the browser for previews, and a connected Vercel account lets users publish a live URL.

  2. Google AI DevelopersOfficialAI score47

    Google launches Gemini 3.5 Transcribe, a speech-to-text model for developers

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that filters out spoken hesitations and accurately grounds technical terms, file names, and code variables against the active context. The model uses visual biasing to incorporate screen-aware context into developer workflows, as demonstrated in Antigravity.

    Video from @googleaidevs's post
  3. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogOfficialAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  4. Andrew NgXAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    AIOpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  5. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon

Aug 21

Aug 21Fri

Aug 20

Aug 20Thu

Aug 19

Aug 19Wed
  1. TinkerOfficialAI score38

    Qwen3.8-27B is now available on Tinker

    AITinker has made Qwen3.8-27B available today. The model is natively multimodal, handling images and video, with flexible thinking control. Tinker says it performs meaningfully better at coding, professional work, research, and long-horizon agentic tasks.

Aug 18

Aug 18Tue
  1. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.