Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. Latent.SpaceXAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  2. Xiaomi MiMoOfficialAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  3. Xiaomi MiMoOfficialAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

    Image from @XiaomiMiMo's post
  4. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  5. Google AI DevelopersOfficialAI score22

    Google demos Gemini 3.8 coding tutor that sees your screen

    AIGoogle AI Developers showed a coding tutor built with Gemini 3.8 Live Extended Thinking that views the user's screen and calls functions to reference the p5.js library. The tutor points out the exact bug on screen and talks the user through the logic.

    Video from @googleaidevs's post
  6. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  7. NVIDIAOfficialAI score38

    Grok 4.7 launches as xAI's most capable model for coding

    AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.

  8. ModelScopeOfficialAI score36

    Qwen Launches RecreationBench for Hybrid Computer-Use Agent Evaluation

    AIQwen introduced RecreationBench, a benchmark of 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass programmatic tests plus VLM-based visual evaluation. The dataset is available on ModelScope.

    Image from @ModelScope2022's post

Sep 20

Sep 20Sun
  1. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. Mike KnoopXAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  2. Thomas DohmkeXAI score24

    Claude Code adds AGENTS.md support in version 2.1.277

    AIClaude Code version 2.1.277 now reads AGENTS.md when a folder has no CLAUDE.md, according to Anthropic engineer Thariq Shihipar's post. The behavior can be toggled in /config. Thomas Dohmke's main post jokingly says AI is finally aligned, with no further technical detail.

  3. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  4. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

Sep 17

Sep 17Thu
  1. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  2. Sherwin WuXAI score44

    ChatGPT for Word launches, bringing ChatGPT and Codex into Microsoft Word

    AIOpenAI has released ChatGPT for Word, letting users access ChatGPT and Codex directly inside Microsoft Word. The post says ChatGPT for Excel and PowerPoint has been growing rapidly, and Word completes that set. The quoted ChatGPT post adds that the tool can draft from notes, rewrite paragraphs, proofread, suggest edits, and flag formatting issues.

  3. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  4. Google AI StudioOfficialAI score58

    Google AI Studio open-sources Speakeasy's OpenAPI SDK generator suite

    AIGoogle AI Studio announced that Speakeasy is open sourcing its OpenAPI client generation suite under AGPLv3, following a May 2026 vendor shutdown that disrupted Google's SDK pipeline. The suite covers SDK generation for 7 languages, an agent-native CLI generator, and a documentation MCP server generator. Google says its pipeline now serves six targets with roughly one engineer maintaining it.

  5. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

  6. Hacker News · Launch HN, YC launches (10+ points)BlogAI score58

    Skillsync launches tool to move AI chat sessions across coding agents

    AISkillsync, a Y Combinator W26 company, launched a tool that converts AI chat sessions between coding agents, including messages, reasoning and tool calls. The conversion engine txcript is open source, and the local-first app runs on Mac with a CLI and MCP support. Only sessions shared into team workspaces leave the user's machine.

  7. JetBrains AI BlogOfficialAI score44

    Building a RAG Pipeline for Semantic Code Search: A Developer Diary

    AIJetBrains describes building Air Context, a RAG pipeline that gives LLM agents semantic code search over real repositories instead of grep. The first installment covers parsing, chunking, and vectorization, arguing that fixed-size line chunks split related code and that structure-aware chunking using language grammar produces better retrieval units.

Sep 16

Sep 16Wed
  1. Amp NewsOfficialAI score50

    Amp Runners Now Serve Multiple Directories and Update Themselves

    AIAmp runners can now serve multiple directories, specified with repeated --dir flags or found automatically with --discover-dirs, which scans Git checkouts up to two levels deep by default. Runners also check for new releases about once an hour, install them, and restart into the new version once no thread is running, at most once every 12 hours, with auto-update disabled via amp.runner.autoUpdate.enabled: false.

  2. Greg BrockmanXAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

  3. Perplexity DevelopersOfficialAI score34

    Perplexity's Search SDK extracts query-relevant passages from URLs for agents

    AIPerplexity says its Search SDK extracts passages relevant to a query from user-provided URLs. Agents can use those passages instead of full pages, keeping unrelated content out of the model context. The company also points to an Agent Skill for installing the Search SDK in coding agents.

  4. TinkerOfficialAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  5. Kilo (acq. by Anaconda)OfficialAI score40

    Kilo Mobile lets users run full AI agent loops from their phone

    AIKilo Mobile now lets users spawn Cloud Agents, start sessions on remote machines, and dictate prompts by voice from a phone. Users can also review and comment on pull requests and approve Security Agent remediations without a laptop. On iPhone, Live Activities show session status on the Lock Screen when an agent needs input.

    Image from @kilocode's post
  6. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  7. Greg BrockmanXAI score22

    Codex voice coming to CarPlay for hands-free road-trip coding

    AIGreg Brockman shared a demo of Codex voice running in CarPlay, letting users build software by speaking while driving. A quoted post from Jonathan Roomer says the setup runs through a third-party iOS app that connects ChatGPT Voice to Codex and his Mac.

Sep 15

Sep 15Tue
  1. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. xAI News (Grok)OfficialAI score43

    Grok Build Adds Memory That Saves Project Notes Between Sessions

    AIGrok Build now has memory, which records conventions, decisions, and project facts after each completed turn and reads them in later sessions. Notes are stored per project plus a global set, and the /dream command organizes them into topic files while /memory opens a read-only browser. The feature is available now and applies to new sessions.