Google reportedly tests Gemini 4 checkpoint "Carbon" matching Opus 5.5 in coding
AIBusiness Insider reports that Google is internally testing a new Gemini 4 checkpoint named Carbon. The checkpoint reportedly matches Opus 5.5 in coding.

Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIBusiness Insider reports that Google is internally testing a new Gemini 4 checkpoint named Carbon. The checkpoint reportedly matches Opus 5.5 in coding.

AILenovo's TianxiCode, paired with DeepSeek-v4.1-Flash, ranked first on the SWE-bench-Live Lite leaderboard with a 71% issue resolution rate and passed official Verified review. The framework combines multi-hop retrieval, autonomous planning with multi-turn tool calling, and test-driven self-correction, and will be applied to Lenovo AI hardware products.
AIAnthropic released Claude Haiku 5.5, its cheapest and fastest small model, which it says costs around 75% less to run than Claude Haiku 4.5. The author argues Anthropic is the only frontier model builder, though the piece also covers Nous Research's $90 million Series B and OpenAI's revenue discrepancy.
AIStanford's CS146S course, taught by Mihail Eric, has published its Week 3 materials on Agent Skills and CLI. The lecture covers how SKILL.md files and scripts encode workflows, and it lists practical advice such as keeping each skill focused, mining one's own transcripts for skill ideas, and writing descriptions that name real trigger phrases.

AISnyk moved its internal support agent, Snyk Assist, into the core Snyk product in September 2026, giving every paying customer access. Built on LangChain and LangGraph with observability in LangSmith, the agent answers questions in plain language and can open support cases or log feature requests. It runs as a single agent behind Slack, web and API surfaces, with tools attached per user permissions.
AIAt the Apsara Conference, Alibaba's Qwen team outlined a roadmap of Qwen4 followed by Qwen4.5 and Qwen5, aiming for 5T to 10T parameters. The article notes that Qwen3.8 reached 2.4T parameters and that Qwen3.8-Flash activates 6B parameters per inference while cutting training cost to one-ninth. It also describes Qwen3.8-Max running model-driven experiments in chip design and inference optimization, and multimodal updates including a video model slated for November.
AIAugment Code is selling select assets, including Cosmos, Auggie CLI, and the Code Context Engine, to Harness, and the product team is moving to Harness. The company says Harness's integrated platform delivers these capabilities to customers more effectively than building them independently. Harness describes itself as building the Autonomous SDLC Platform for shipping AI-written code across enterprises.
Why it matters: The announcement shows how a coding AI company is folding its products into a larger software delivery platform, a shift that shapes how enterprise teams will buy these tools.
AILegalOn cut its estimated daily Codex costs by 65% while maintaining development speed. The company matched Astra, Sol, and Luna to tasks and managed budgets strategically.
AIA suspected Chinese-speaking attacker breached several South Korean financial institutions between late September and early October 2026, reportedly stealing over 25,000 records from Shinhan Bank alone. The attacker used ARTEX, a Chinese open-source tool that uses AI language models to automate finding security flaws, and models named in the report include DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6.
AIMicrosoft has reportedly cut monthly AI spending limits for employees from $100,000 to about $10,000 and told staff to use tools such as GitHub Copilot. Meta's Claude Code users reportedly fell from about 60,000 to 30,000 as it shifts to its own MetaCode and Muse Code. The source says customer access to Claude through Microsoft's platforms continues.
AIGitHub reports that Git events on the platform rose from 218.2 billion to 473.3 billion per month between September 2025 and August 2026. It says agent workloads push write throughput and merge contention beyond what its current replica-based architecture handles well, so it is separating durable storage from compute while GitHub keeps running. The article states internal benchmarks reached up to 35 times higher write throughput.
Why it matters: The post links rising Git event volume to specific architectural bottlenecks, showing why agent workloads strain write paths and how GitHub plans to separate storage from compute.
AIAt Runtime, Modal (@modal) shared how typesafeai's Jev hit a trillion tokens a day by its first weekend. The post says he also argued that benchmarks hurt the industry and that a SaaSapalooza is beginning as the "software is over" era ends.
AIAnthropic's ClaudeDevs account says Pro or Max plan subscribers on September 23 can still claim a one-time bonus credit for cloud sessions by running /claim-credit in Claude Code by October 7 at 11:59pm PT. Cloud sessions draw on this credit first before counting toward plan limits, and the credit expires November 4.
AIAccording to Gergely Orosz, GitHub now records more AI-generated pull requests per month than human-written ones. He says the volume of AI-generated PRs is about three times what it was in December 2023.
AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.
AIOpenAI announced more than 20 updates at DevDay 2026, including dots always-on agents, GPT-6.1 Sol, Ultrafast token generation, ChatGPT Space, and a $500/month Pro 500 plan. GPT-6.1 Sol is priced at $2 input and $10 output per 1M tokens and is available in the API as gpt-6.1-sol. Ultrafast generates tokens up to 8x faster in Codex and up to 6x faster in the API.
Why it matters: The post lists dozens of OpenAI DevDay 2026 changes across models, agents, plans, and APIs, useful for scanning what shipped and who gets access.

AITogether AI ranks first on OpenRouter token share among top open coding models, with Z.ai's GLM 5.3 Flash at 29.2%, DeepSeek's V4.1 Flash at 25.6%, and Moonshot's Kimi K3 at 18.9%. The post positions Together as a go-to provider for running coding agents on open models.

AIProximal, a year-old startup supplying coding data to AI labs, raised funding from General Catalyst at a $300m valuation. In the past 10 months, it has surpassed $200m in annualized revenue, reflecting the labs' ongoing demand for data.
AIMatei Zaharia said autoresearch produced strong results that are being integrated into a model serving stack. The post gives no specific figures, benchmarks, or product names. Background context from a related post says Databricks ranked #1 on NVIDIA's SOL-ExecBench kernel leaderboard across all four tracks using agents.
AIGoogle and Kaggle launched the Gemma 4 Developer Agent Competition, challenging developers to build coding agents that run offline on consumer hardware rather than relying on cloud API models. The total prize pool exceeds $110,000, with entries due November 2, 2026, and a starter kit is available on Kaggle.

AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.
Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.
AIWe added planning mode in 2025 and deleted it from the product earlier this year. Users wanted a way to explicitly plan with the model so we added this opt in slash command. Understood that the timing couldn’t be worse since it appears like we’re adding this for the first time. Have a great weekend folks, lots more to come in the coming weeks!
AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.
AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.
AIDatabricks rolled out Astra to all of its engineers and reported internal measurements. Engineers given Astra increased overall coding spend by around 60% compared with baseline, and Astra outperformed prior top models on highly complex system design tasks, while the gain on medium or low complexity tasks was unclear.
AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.
Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.
AIFactory has raised $200M at a $5B valuation from investors including Blackstone, Khosla Ventures, and Sequoia Capital, bringing its total funding to over $400 million. The company says it will use the capital to accelerate research, product, and global go-to-market efforts. Factory says hundreds of thousands of developers use its platform, with customers including Nvidia, Blackstone, and T-Mobile.
AICognition has raised over $2B at a $48B valuation, led by Andreessen Horowitz and Accel. The company says run-rate revenue grew from $492M to almost $900M since its May round, and Devin now offers Auto-Triage, Security Swarm, and Automations.
AIMihail Eric announced the 2026 edition of his Stanford course "The Modern Software Developer," with 85% of the Fall 2025 material replaced by AI-native topics such as agent skills, context engineering, and agentic code review. Students will ship pull requests to real open-source AI repositories, with partners including Browserbase, HeyGen, and CopilotKit offering mentorship.
AICursor's CEO Michael Truell says OpenAI announced plans to block Cursor users from accessing OpenAI models in three months. He states that OpenAI models account for about 5% of Cursor user traffic and that Cursor is speaking with OpenAI to resolve the issue.
AICursor's acquisition by SpaceX has officially closed, and Cursor will join the SpaceXAI team. The stated goal is to help make Grok the world's most useful AI and to improve Grok Build, Grok Bot, Grok API, Cursor, and more.
Why it matters: The acquisition closing ties Cursor's coding tools to Grok's product line, which changes how the two products may be developed and sold together.
AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.
AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.
AISebastien Bubeck says GPT, using Codex, found a counterexample to a math problem, with its reasoning described as superb. He highlights that Codex's write-up undermines the "illusion of thinking" argument and shares the exact prompt used on internal Codex.
AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.
AIXiaomi MiMo's account celebrated growing developer adoption of its open-weights models, crediting Cline for building on MiMo. Cline's linked post announced a $9.99/month subscription offering 2-5x discounted access to GLM-5.2 and other open-weight models including DeepSeek, Kimi, MiniMax, MiMo, and Qwen, with a $1.99 promo for sign-ups via npm i -g cline.
AISpaceX has exercised its option to acquire Cursor in an all-stock transaction, aiming to build highly useful AI models. The quoted announcement says SpaceXAI has spent several months jointly training a model with Cursor, to be released soon in Cursor and Grok Build.
AISpaceX has exercised its option to acquire Cursor in an all-stock transaction, with the stated goal of building the world's most useful AI models. Over the past few months, SpaceXAI has been jointly training a model with Cursor, which will be released in Cursor and Grok Build soon.
AICognition introduced the AI Productivity Guarantee, under which it will issue credits up to $10M if Devin delivers less engineering value than enterprise customers pay for. The company uses an AI estimator to measure hours of productive output, validated against engineers' own estimates of how long the same work would have taken by hand. Value is converted to dollars at a standard global rate and compared against each customer's consumption near the end of the annual contract.
Why it matters: The post explains how Cognition estimates Devin's output in hours and backs the estimate with a $10M credit commitment, a concrete model for measuring AI vendor value.
AIXiaomi MiMo will credit a free Token Plan of the same tier to users whose plan expired before May 27. Users who renewed after May 27 will receive a balance matching their renewal amount, usable toward API usage, and both are valid for one calendar month.
