Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

May 8

May 8Fri
  1. PaddlePaddleOfficialAI score12

    ERNIE 5.1 ranks #4 in Search Arena, official release soon

    AI🎉 Quoted teaser exception applies: the main post is only a reaction, so the quoted post carries the news. ERNIE 5.1 is now ranked #4 in Search Arena, making it the only Chinese model in the Top 10. An official release is coming very soon.

  2. Jan LeikeXAI score14

    Jan Leike announces a new research project at Anthropic

    AIJan Leike says he is starting a new research project at Anthropic and calls it very exciting. He adds that making AGI go well requires many things, with alignment being only one of them, and promises more details soon.

May 6

May 6Wed
  1. OpenAI Alignment Research BlogOfficialAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    AIOpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    Why it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

  2. Eugene YanXAI score44

    Anthropic gains more compute via SpaceX, boosting Claude rate limits

    AIAnthropic announced a partnership with SpaceX that will give it more compute, according to the post. The related Claude post says Claude Code's 5-hour rate limits are doubled for Pro, Max, and Team plans, peak-hour reductions are removed for Pro and Max, and Opus API rate limits are substantially raised.

May 2

May 2Sat
  1. Nick TurleyXAI score27

    ChatGPT's new image feature sees usage up over 50% in weeks

    AIOpenAI's Nick Turley reports that usage of the new ChatGPT images feature rose more than 50% within a few weeks. Nearly 60% of daily users are newly logged-in users, and the feature is being used across home design, learning, work graphics, and creative projects.

May 1

May 1Fri
  1. ReflectionOfficialAI score22

    Reflection says shared Pentagon understanding sets AI deployment precedent

    AIReflection says its mission is to build open, safe, and accessible frontier intelligence, and that it has a responsibility to shape how AI is deployed. The company says a shared understanding with the Pentagon sets a precedent for how AI labs could work across the U.S. government, from supporting service members to scientists.

  2. ReflectionOfficialAI score38

    Reflection joins AI coalition on responsible U.S. government deployment

    AIReflection has joined a coalition including AWS, Microsoft, OpenAI, Google, and Nvidia on a framework governing how the U.S. government licenses and deploys AI. The agreement, which includes a non-binding memorandum of understanding with the DoW, commits to safety, red-teaming, and ongoing evaluation and explicitly prohibits unlawful mass surveillance and autonomous weapon use. Reflection says it will keep its commitment to open source while customizing its models for scientists in national labs.

Apr 30

Apr 30Thu
  1. Mark ChenXAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

    Why it matters: The post links a single cyber-range eval to OpenAI's own safety framing, so readers can weigh the result against the company's stated risk and deployment position.

  2. koray kavukcuogluXAI score12

    Alphabet named to TIME's 2026 Most Influential Companies list

    AIAlphabet has been recognized in TIME's 2026 Most Influential Companies list, with Google's Koray Kavukcuoglu crediting teams for turning AI breakthroughs into everyday user features. TIME's cover story says Sundar Pichai's 2016 "AI-first" push, including custom chips, Cloud, YouTube, and deep AI research, has paid off.

Apr 29

Apr 29Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score34

    Cognition opens Singapore headquarters for Asia-Pacific push with Devin

    AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

  2. Andy JassyXAI score47

    OpenAI models coming to Amazon Bedrock in coming weeks

    AIAndy Jassy announced that OpenAI's models will become available directly to customers on Amazon Bedrock in the coming weeks. The availability will be alongside the upcoming Stateful Runtime Environment, giving builders more model choice. More details are expected at the AWS event in San Francisco tomorrow.

Apr 24

Apr 24Fri
  1. Andy JassyXAI score42

    Meta commits to tens of millions of AWS Graviton cores

    AIMeta has committed to tens of millions of AWS Graviton cores, according to Andy Jassy, Amazon's CEO. Jassy argues agentic AI is becoming as much a CPU story as a GPU story, since multi-step orchestration and real-time reasoning are CPU-intensive. He says the Graviton5 instances deliver up to 33% lower latency between cores.

    Video from @ajassy's post

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed
  1. Anthropic EngineeringOfficialAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

Apr 21

Apr 21Tue
  1. Michael TruellXAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

    Why it matters: The post shows a partnership with conditional acquisition terms, which matters for judging how Cursor's coding products and model training may be developed.

Apr 18

Apr 18Sat
  1. Mark ChenXAI score42

    Mark Chen says science remains vital, names Bubeck and Helkky to lead

    AIOpenAI's Mark Chen rejected the claim that science is no longer a priority, saying AI should accelerate discovery in partnership with scientists. He said Sébastien Bubeck and Helkky are taking on this mandate, which the quoted post frames as a response to reports that OpenAI's VP of Science is leaving.

Apr 17

Apr 17Fri
  1. NVIDIA AI DeveloperOfficialAI score8

    NVIDIA credits Picomoorhead for OpenClaw contribution

    AINVIDIA AI Developer thanked @Picomoorhead for a contribution to OpenClaw, which the post links to via @openclaw. The post provides no details about what the contribution was or what it enables.

  2. NVIDIA AI DeveloperOfficialAI score23

    NVIDIA's Cosmos Cookoff winners showcase Cosmos Reason 2 projects

    AINVIDIA highlighted the developers who won the Cosmos Cookoff and how they used Cosmos Reason 2 to build projects spanning disaster-response drones, explainable visual AI, and intelligent security systems. The post links to a YouTube showcase and a LinkedIn recap of the event.

    Image from @NVIDIAAIDev's post

Apr 13

Apr 13Mon
  1. Mira MuratiXAI score26

    Thinking Machines welcomes Workshop Labs founders Luke Drago and @LRudL_

    AIThinking Machines has welcomed Luke Drago and @LRudL_, who co-founded Workshop Labs to build AI that keeps humans relevant. Mira Murati says they will continue that mission at Thinking Machines, which builds powerful AI systems that think alongside humans and extend human agency. She ties the hire to the company's broader work, including Tinker, research grants, and advancing the frontier.

Apr 8

Apr 8Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score31

    Cognition Expands to Japan, Appoints Takumi Masai to Lead Devin Launch

    AICognition is expanding into Japan, its first step into Asia, and has appointed Takumi Masai as Japan President and General Manager to lead a local team working with Japanese enterprises. DeNA has used Devin to more than double operational efficiency across multiple engineering functions, and Mizuho Securities has deployed it as one of the first large-scale financial institutions in Japan.

Apr 7

Apr 7Tue
  1. Dario AmodeiXAI score62

    Anthropic's Dario Amodei says new Mythos Preview model shows a large jump in cyber capabilities

    AIDario Amodei says the company has tracked growing cyber capabilities in AI models for years, which arise from their general coding proficiency. He states that the new model, Mythos Preview, represents a particularly large step up in those capabilities.

    Why it matters: The post links rising cyber capability to general coding skill, and names a notable jump in a new model, Mythos Preview.

  2. Dario AmodeiXAI score72

    Dario Amodei backs Project Glasswing to counter AI-driven cyber threats

    AIDario Amodei said many of the world's leading companies have joined Project Glasswing, an effort to address cyber threats posed by increasingly capable AI systems. The initiative was introduced by Anthropic and is powered by its newest frontier model, Claude Mythos Preview, which the quoted post says can find software vulnerabilities better than all but the most skilled humans.

    Why it matters: The post gives a concrete example of how a frontier AI lab is organizing industry partners around AI-driven software vulnerability discovery.

  3. Andy JassyXAI score36

    Uber uses AWS Graviton4 and Trainium3 chips for rides and AI

    AIUber is running its ride and delivery matching on AWS Graviton4 chips and training its AI models on Trainium3, according to Amazon CEO Andy Jassy. Jassy says the Graviton4 setup matches riders with drivers in fractions of a second at lower cost, while Trainium3 helps make rides smarter over time.

    Image from @ajassy's post

Apr 6

Apr 6Mon
  1. OpenAI Alignment Research BlogOfficialAI score31

    OpenAI opens applications for Safety Fellowship on AI safety and alignment research

    AIOpenAI announced applications for its Safety Fellowship, a pilot program supporting external researchers, engineers, and practitioners in safety and alignment research on advanced AI systems. The program runs from September 14, 2026 through February 5, 2027, with a monthly stipend, compute support, API credits, and mentorship, and fellows are expected to produce a substantial output such as a paper, benchmark, or dataset. Applications close May 3, and successful applicants will be notified by July 25.

Mar 30

Mar 30Mon

Mar 27

Mar 27Fri

Mar 26

Mar 26Thu

Mar 20

Mar 20Fri
  1. Aman SangerXAI score55

    Cursor's Composer 2 is built on Kimi k2.5 base model with added training

    AIAman Sanger says Cursor's team evaluated many base models on perplexity-based evals and found Kimi k2.5 the strongest. Composer 2 was then built with continued pretraining and a 4x scale-up of high-compute RL, with Fireworks providing inference and RL samplers. The author admits Cursor should have named the Kimi base in its launch blog and says it will do so for the next model.

  2. Aman SangerXAI score22

    Cursor's Composer 2 model praised, built on an open-source base

    AIAman Sanger of Cursor says Composer 2 is a really good model and he is excited for more people to try it. The quoted reply from Lee Robinson says Composer 2 started from an open-source base, with only about one-quarter of the final model's compute coming from that base. Cursor plans full pretraining in the future and says it is following the license through its inference partner terms.

Mar 16

Mar 16Mon

Mar 10

Mar 10Tue