OpenAI announced more than 20 updates at DevDay 2026, including always-on agents on GPT-6 Astra, GPT-6.1 Sol priced at a fifth of Astra's API cost, and a new $500/month Pro 500 plan.
Anthropic launched Claude Sonnet 5.5 at $2/$10 per million tokens, 30%+ faster than Sonnet 5, with thinking always on. Reuters reported the FTC is probing OpenAI, Anthropic and other labs over rogue AI agents.
https://x.com/i/article/2106699200201179136
OpenAI and Anthropic weekly: DevDay 2026, Sonnet 5.5, FTC probe, Australia apology (Week 40, 2026)
OpenAI announced more than 20 updates at DevDay 2026 including dots, always-on agents on GPT-6 Astra with their own cloud computer and browser rolling out to Pro and Business Premium in eligible markets and as an off-by-default beta for Enterprise, GPT-6.1 Sol at a fifth of Astra's price in the API, ChatGPT Work and Codex but not yet in Chat, an Ultrafast tier with up to 8x faster token generation in Codex, a new $500/month Pro 500 plan as the only Pro tier with Ultrafast, Pro 200 reopened at a lower 10x Plus allowance with existing subscribers keeping 20x through October 29 and also getting additional credits, ChatGPT Space replacing Library with pages you edit with teammates, a Meetings plugin beta on the macOS app and @ ChatGPT in Slack, published a model guide for the GPT-6 family with multi-agent delegation in beta for GPT-6.1 Sol, fixed an image encoding bug that was degrading image understanding in GPT-6 Sol and GPT-6 Luna, added virtual try-on and a Scan mode to ChatGPT and brought Finances to Free and Go users in the US, announced GPT-Synopsys, a chip design model built with Synopsys under a revenue-sharing agreement, doubled its Lenfest support with a new $5 million commitment plus up to $5 million in credits, partnered with America's SBDC on small business AI training, outlined structured safety documentation it says should be required before continuing any frontier RL training run, ideally rising to full safety cases, disrupted a coordinated model-distillation campaign whose core cluster it attributes to individuals linked to Moonshot AI, and apologized to Australia after an experimental internal-only model gained unauthorized non-public access to government systems in June and retrieved internal files and credentials, admitting it should have shared preliminary findings sooner, pausing tool-use training for its most capable models and sending Jason Kwon to a Sydney committee hearing on October 6, as AP reported Trump saying OpenAI, Anthropic and other top labs signed a voluntary self-policing accord at the White House and Sam Altman told CNBC the company is pacing its progress, which includes sometimes not training a model, Reuters reported the FTC is probing OpenAI, Anthropic and other labs over rogue AI agents with compelled executive testimony planned, the WSJ reported OpenAI parted ways with three safety researchers for allegedly sharing confidential information days after scrapping the GPT-6.1 Astra launch over safety concerns, a departing Preparedness Framework lead wrote in The Atlantic that the sprint-driven culture is not careful enough, and The Information reported the hire of former White House AI policy lead Thomas Lind for its national security policy team
Anthropic launched Claude Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5 at the same $2/$10 pricing as Sonnet 5 but 30%+ faster and up to 30% cheaper per task, with thinking always on, the first Sonnet with cyber safeguards that fall back to Sonnet 5 on higher-risk tasks, and Haiku 5.5 due in the coming weeks, made Claude for Government generally available through a FedRAMP High environment with no seat fees and prepaid usage, paired Claude Managed Agents with NVIDIA's open source OpenShell so agents never see their credentials and every action is blocked unless a rule allows it, introduced mods for Claude Code, TypeScript functions in plugins that rewrite prompts, replace UI and swap out built-in features, plus a built-in You should know mod that flags important information you might miss, added eval building and hillclimbing commands to the claude-api skill with an internal example where Sonnet 5 at low effort beat an Opus 4.8 setup at about a fifth of the cost, launched Claude Frontier Academy with $100 million to train 10,000 Frontier Deployed Engineers by the end of 2027, is running a two-week promotion where starting with Claude Design, Slides or Docs makes the rest of that conversation use 50% less of your usage limits through October 15, opened a new Anthropic Interviewer study on what you want from AI through October 6 with an option to publish your full interview that it asks you to treat as permanent, found Zhipu AI's open-weight GLM-5.3 close to Claude Mythos Preview at building end-to-end cyber exploits but released without meaningful safeguards and about four months behind the US frontier, published a robot exposure index finding robots can already do 74% of physical tasks in the US but are cost-competitive for just 0.3% of all work, and a Science Blog guest post by Harvard physicist Matthew Schwartz, a visiting researcher at Anthropic, on BootLoops, his open-source toolkit built with Claude Fable 5 that is not an Anthropic project, and more
OpenAI fixed an image encoding bug in GPT-6 Sol and GPT-6 Luna that was degrading image understanding, with better results now on visual tasks in the API and Codex, including computer use, and a recommendation to rerun evals if your workflows use image inputs
OpenAI outlined safety cases for frontier RL training runs as structured safety documentation it believes should be required before continuing any such run, with technical guidelines across alignment training, containment, and monitoring, plus operational practices like dissents, senior leadership vetoes, fail-closed auto-pausing, audits, and a misalignment on-call that can page executives including the CEO, all described as current recommendations being implemented at OpenAI
Anthropic added eval design and hillclimbing guidance to the claude-api skill for Claude Code, where /claude-api build-eval interviews you and builds an eval in your codebase from production traffic or synthetic data with a validated grader and baseline, and /claude-api hillclimb tries one change per round against a held-out test split and reverts anything that only helps the train set, with an internal customer support example ending on Sonnet 5 at low effort scoring 90.5% vs 78.6% for the original Opus 4.8 setup on held-out tickets at about one fifth of the cost
Anthropic launched Claude Sonnet 5.5, the second Claude 5.5 model as a faster, lower-cost complement to Opus 5.5, available now on all platforms at the same $2/$10 per million tokens as Sonnet 5 but 30%+ faster and up to 30% cheaper per task, with thinking always on and a new between_tools setting replacing disabled, the first Sonnet with cyber safeguards that fall back to Sonnet 5 on higher-risk tasks and account-bound preserved thinking against distillation, and Haiku 5.5 due in the coming weeks
OpenAI apologized to Australia for unauthorized model access to government websites during internal training and evaluation in June, found in a mid-August review after the Hugging Face incident, where an experimental internal-only model gained non-public access to Services Australia's Medicare Statistics Reporting Service and retrieved internal files and credentials, with further activity at NSW BOCSAR, the Victorian Department of Health, and AIHW but no individual records accessed, admitting it should have shared preliminary findings sooner, noting it has paused tool-use training for its most capable models, and committing to dedicated support, Daybreak fund credits and an Australian taskforce due by year end, with Chief Strategy Officer Jason Kwon appearing before the Joint Select Committee on AI in Sydney on Tuesday, October 6
OpenAI doubled its support for the Lenfest AI Collaborative and Fellowship Program with a new $5 million commitment plus up to $5 million in software credits and engineering support, after two years of embedding AI engineering fellows in 11 US news organizations, with the Lenfest Institute inviting a new cohort of news organizations and planning to turn the strongest fellowship projects into reusable tools, plugins, and playbooks
Anthropic is pairing Claude Managed Agents with NVIDIA's open source OpenShell under NVIDIA's new Open Agent Safety Platform, with Managed Agents holding an agent's credentials in a vault it never sees and OpenShell blocking every file, network and tool action unless a rule allows it, logging each decision and using a policy prover to confirm what the agent can reach, Managed Agents available today in a sandbox you control on your own infrastructure or with a managed provider, OpenShell on GitHub under Apache 2.0, and Notion, Rakuten and Asana named as Managed Agents users
Anthropic launched a new study run by Anthropic Interviewer on your experiences with AI and what you want from the companies building it, open September 29 to October 6 to Free, Pro and Max users on Claude and Claude Code with accounts at least two weeks old, taking about 15 minutes per interview, to shape the Anthropic Institute's research, following last December's study with 81,000 people, and this time you can optionally publish your full interview with only your country attached, a choice Anthropic asks you to treat as permanent since copies can't be recalled once saved
OpenAI announced more than 20 updates at DevDay 2026 including dots, always-on agents on GPT-6 Astra with their own cloud computer and browser rolling out to Pro and Business Premium in eligible markets and as an off-by-default beta for Enterprise, GPT-6.1 Sol at a fifth of Astra's price ($2 input, $10 output, $0.10 cached input per 1M tokens) in the API, ChatGPT Work and Codex but not yet in Chat, the Ultrafast tier with up to 8x faster token generation in Codex for Astra today and GPT-6.1 Sol in the coming days, a new $500/month Pro 500 plan as the only Pro tier with Ultrafast, Pro 200 reopened at a lower 10x Plus allowance, with existing subscribers keeping 20x through October 29 and getting additional credits, ChatGPT Space replacing Library with pages you edit together with teammates, ChatGPT and your dot, the Meetings plugin in beta on the macOS desktop app, @ ChatGPT in Slack and Teams for Business and Enterprise, your ChatGPT plan allowance usable in 16 partner tools including Devin, Notion, Vercel, T3, OpenClaw and Dactyl, OpenAI Marketplace letting eligible enterprises put part of their OpenAI commitment toward 32 partner tools, Private Intelligence with per-project Zero Data Retention policies, reusable Codex cloud environments, a refreshed Codex CLI with /agents, /fork and /voice, Codex Security Cloud in research preview with Daybreak Blue models, a Decisions API on Luna in limited preview, Agents API updates with computer use in one API call and Bedrock Managed Agents on AWS, plugin extensions, Sign in with ChatGPT and 1.2B weekly users
Anthropic found Zhipu AI's open-weight GLM-5.3 close to Claude Mythos Preview at autonomously building end-to-end cyber exploits (50 vs 56 of 410 attempts on ExploitBench) but released without meaningful safeguards, with simple techniques bypassing its refusals 64% to 100% of the time in simulated tests but failing against safeguarded Claude models, an abliterated copy costing roughly $4,400 in compute cutting refusals from above 90% to as low as 2%, a day-long session finding several previously unknown vulnerabilities in a popular browser's JavaScript engine and chaining them into a working exploit, now disclosed to the maintainer, and the smaller GLM-5.3-Flash chaining a Chrome exploit for $20.40 in API costs, broadly matching NIST CAISI's assessment of it as the most cyber-capable open-weight model to date, about four months behind the US frontier
Trump says top AI firms signed a voluntary self-policing accord at the White House, with Anthropic CEO Dario Amodei, OpenAI President Greg Brockman, Google's Sundar Pichai, Meta's Mark Zuckerberg, Nvidia's Jensen Huang and xAI's Elon Musk committing to robust internal controls, an independent external auditor and a board committee to review audit reports, leaving the door open to codifying the steps into law later, Trump calling it "morally binding", Amodei saying the mechanism for addressing AI's "very real risks" is still under discussion, and Sam Altman telling CNBC that OpenAI is "pacing our progress, which includes sometimes not training a model" after AP reports it halted a new model's rollout a day earlier over safety concerns
OpenAI partnered with America's SBDC on AI training for small businesses planning to train around 150 advisors through OpenAI Academy and reach at least 1,000 businesses with hands-on ChatGPT workshops, and published a Small Businesses, Bigger Capabilities report finding roughly 4 million employees at companies under 500 people used its tools in one September week, with agentic output tokens for ChatGPT Work and Codex making up two-thirds of small business output tokens in August, up from a third in April
OpenAI disrupted a coordinated campaign to extract protected reasoning from its models that began on July 1, peaked with 16,000 requests from over 4,000 users on July 24 and 25, spanned a cluster of more than 15,000 users fully disrupted by July 28, used tricks like asking a model in another conversation to decrypt copied encrypted reasoning, and is attributed in its core cluster to individuals associated with Moonshot AI, the developer of Kimi, with findings shared through the Frontier Model Forum
Anthropic published a robot exposure index built with Claude scoring O*NET job tasks by the environment a robot needs, finding robots can already do 74% of physical tasks in the US covering 34% of working hours but mostly in controlled settings, are cost-competitive for just 0.3% of tasks with 40 years needed to reach 10% at past price trends, and that about 80% of work time is exposed to robots or LLMs combined, with driving and warehouse jobs most exposed and nursing and repair jobs barely at all
Claude for Government is now generally available for federal and state agencies through a FedRAMP High authorized environment after a public beta since July, with no seat fees and prepaid usage under a hard not-to-exceed cap, department-level admins allocating spend and model limits per tier, audit logs and two-person approval for sensitive operations, conversation history kept on the agency-managed device, and Claude Code CLI and Claude for Microsoft 365 rolling out in early access through the same environment
The Federal Trade Commission is probing Anthropic, OpenAI, and other AI labs over dangers their technology poses to consumers, according to Reuters, in the first official US enforcement action on rogue AI agents, with formal information demands and compelled executive testimony planned, including from METR, after incidents like OpenAI agents probing Hugging Face for vulnerabilities before a large-scale attack, Chairman Andrew Ferguson suggesting developers whose cybersecurity-testing agents cause hacks should be liable, and Anthropic's IPO prospectus flagging agentic AI as a significant and unpredictable legal risk
OpenAI and Synopsys announced GPT-Synopsys, a specialized chip design model under a multi-year agreement with revenue sharing and joint go-to-market, built with licensed Synopsys electronic design automation tools so the model runs them like an expert engineer and iterates on power, performance, and area optimization and verification closure, running on OpenAI-hosted infrastructure integrated with Synopsys.ai and Synopsys Autopilot, with early engagements underway with semiconductor customers and customer data not used for training
Anthropic published a Science Blog guest post by Harvard physicist Matthew Schwartz on the impedance mismatch between how scientists want to work with LLMs and what they do well, describing BootLoops, his open-source toolkit for exact calculations in quantitative science built with Claude Fable 5 around the semi-numerical S-matrix bootstrap, which reproduced his paper's results in about 20 minutes in the course of porting his code, computed elliptic Feynman integrals, and surfaced connections to ecology, population genetics, and a dozen other fields that he steered with domain experts, with the disclosure that Schwartz is a visiting researcher at Anthropic and BootLoops is not an Anthropic project
Anthropic introduced mods for Claude Code small TypeScript functions shipped inside plugins that can rewrite prompts, add or replace UI, and swap out built-in features like /diff, with sample mods such as Blast Radius for catching risky shell commands before they run and Replay Theater for stepping through file edits, a sec-default mod loading first on Team and Enterprise plans, and a warning that mods run unsandboxed with the same machine access as Claude Code
Anthropic is running a two-week promotion in the Claude app where creating a design, deck, or doc with Claude Design, Slides, or Docs makes the work that follows in that conversation use 50% less of your usage limits, applied automatically on Pro, Max, and Team plans through October 15, with Claude Sonnet 5.5 suggested for slides that need minimal editing
ChatGPT added virtual try-on for clothing and accessories with a Try on button on product listings that generates the image from a selfie saved as a reference photo you can manage in Settings, saving products to Favorites and folders in the Library, and a new Scan mode in the camera on iOS that combines multiple pages into a single PDF
OpenAI parted ways with three researchers for allegedly sharing confidential information with a third-party AI safety organization, according to the WSJ, whose sources say the three worked on its safety team, with OpenAI confirming in a statement that it "parted ways with three individuals for violating our policies on accessing and handling sensitive company information", days after scrapping the GPT-6.1 Astra launch over safety concerns, and a departing OpenAI employee who led the Preparedness Framework and safety reports for 12 launches wrote in The Atlantic that the company's sprint-driven culture is not careful enough
Anthropic added a You should know plugin to Claude Code a built-in mod that spins off a sideagent to observe Claude's output and flag important information you might otherwise miss, enabled via /plugin enable cc-plugin-you-should-know@builtin, with mods explained in a new blog post
OpenAI expanded Finances in ChatGPT to Free and Go users in the US on web, iOS, and Android, letting you connect accounts through Plaid and Experian to find forgotten subscriptions, spot unfamiliar charges, build a budget from actual spending, track your credit score, plan debt payoff, and review your investment mix
OpenAI published a model guide for the GPT-6 family covering when to pick GPT-6 Astra, GPT-6.1 Sol, or GPT-6 Luna, how to tune reasoning effort and Fast or Ultrafast speed, cutting cost with prompt caching and compaction, and keeping long-running work on track with mid-turn steering, async tool calling, and multi-agent delegation, which is currently in beta for GPT-6.1 Sol
Anthropic launched Claude Frontier Academy with a $100 million commitment to train 10,000 Frontier Deployed Engineers by the end of 2027, starting with a multi-day in-person program and a graded practical, then a 12-week residency leading a real Claude use case at their own organization, by nomination only, with first cohorts in San Francisco, New York, and London drawn from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley, and Novo Nordisk, and the first Claude Frontier Deployed Engineer badges expected in early 2027
OpenAI hired former White House AI policy lead Thomas Lind to lead cyber and strategic risk on its national security policy team under Sasha Baker, according to The Information, after Lind headed AI policy at the Office of the National Cyber Director, which shaped the June executive order setting up voluntary government review of models before release, following OpenAI's July hire of Dean Ball from the Office of Science and Technology Policy
That's the week done, more next weekend. The bit I keep thinking about from Anthropic's new eval and hillclimbing skill for Claude Code is the discipline in it: change one thing per round, keep it only if it still holds on tickets the model hasn't seen, revert the rest. I run a low-tech version of that on my own prompts, tweak one line, try it on the next real task, keep what survives. Those survivors live as templates in AIPRM, our browser extension with thousands of ready-made prompts, plus tones and writing styles you can apply in ChatGPT and Claude, so a prompt that worked once keeps working without retyping it every time. Have a look at AIPRM.com.
Source: Tibor Blaho · x.comPublished · added here