Skip to contentSkip to stories

Updated

#Industry news

May 26

May 26Tue
  1. Andy JassyAI score5

    Jassy and Amazon teammates volunteer at Seattle's Community Lunch

    AIAmazon CEO Andy Jassy joined Amazon teammates to prep meals at Community Lunch on Capitol Hill in Seattle during Global Month of Volunteering in May. The nonprofit serves free, made-from-scratch meals to 200-300 low-income and unhoused neighbors each weekday, and nearly 2,000 Amazonians have volunteered there since 2023. Amazonians completed more than 500,000 volunteer activities across 55 countries last year.

May 24

May 24Sun
  1. Benedict EvansAI score42

    Benedict Evans: Predicting which jobs AI will expose is largely impossible

    AIBenedict Evans argues that predicting AI's job impact by occupation is mostly impossible, citing accountants: despite a century of accounting automation from calculators to spreadsheets and ERP systems, the number of accountants kept rising. He says job titles and business models change over time, so exposure scores based on census categories can mislead.

May 22

May 22Fri

May 21

May 21Thu
  1. Mark ChenAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

May 19

May 19Tue

May 18

May 18Mon

May 16

May 16Sat

May 15

May 15Fri

May 14

May 14Thu

May 12

May 12Tue

May 11

May 11Mon

May 8

May 8Fri

May 6

May 6Wed
  1. OpenAI Alignment Research BlogAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    AIOpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    Why it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

May 2

May 2Sat

Apr 30

Apr 30Thu
  1. Mark ChenAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)AI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

  2. Soumith ChintalaAI score15

    This is kinda interesting.

    AIAnthro probably needs to scale up Account Support via Claude or via humans (traditional account managers). Would be funny if they chose humans. Alternatively, more and more companies can probably go multi-AI with open harnesses. Similar problems to the cloud era enterprise pains. I'm sure these issues would apply to all the other AI providers.

Apr 24

Apr 24Fri

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed
  1. Anthropic EngineeringAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

Apr 21

Apr 21Tue
  1. Michael TruellAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

Apr 18

Apr 18Sat

Apr 17

Apr 17Fri

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.