Skip to contentSkip to stories

Updated

AI safety

Showing low-relevance items too. Hide low-relevance items

Jun 10

Jun 10Wed
  1. Jeremy HowardXAI score12

    Jeremy Howard jokingly accuses Anthropic of sabotaging AI research

    AIJeremy Howard posts a sarcastic jab suggesting an unnamed AI developer is silently sabotaging experiments to slow scientific progress and protect its technology lead. He pairs the remark with the mocking nickname "Sophanthropic," a play on Anthropic, and the post names no specific evidence.

    Image from @jeremyphoward's post

Jun 3

Jun 3Wed
  1. Mark ChenXAI score25

    Mark Chen says OpenAI's models could match Mythos on cyber vulnerabilities

    AIOpenAI's Mark Chen said that after Mythos showed AI models can prove 80-year-old theorems, he expected them to also find cyber vulnerabilities, and they did. He added that researchers in math may now be thinking the same idea in reverse, applying cybersecurity-style capability to mathematics. The post offers no specific models, benchmarks, or figures.

Jun 2

Jun 2Tue

May 28

May 28Thu
  1. Sam BowmanXAI score38

    Anthropic highlights AI for transparency in Claude Opus 4.8 system card

    AISam Bowman says he is excited about alignment assessments in the recent system card for Claude Opus 4.8, crediting @MaskedTorah. He argues AI systems have considerable underexplored potential for transparency and coordination. The quoted Claude announcement describes Opus 4.8 as improving on Opus 4.7 with sharper judgment and more honest self-assessment of progress.

    Image from @sleepinyourhat's post

May 25

May 25Mon
  1. Chris OlahXAI score44

    Dario Amodei speaks at Vatican presentation of Magnifica Humanitas on AI

    AIAnthropic co-founder Dario Amodei spoke at the Vatican's presentation of Magnifica Humanitas, arguing that AI's questions extend beyond the AI research community. He said frontier labs face commercial, geopolitical, and competitive incentives that can conflict with doing the right thing, so outside voices from religion, civil society, academia, and government are needed. He described AI models as grown rather than engineered, and framed three questions for the Church's discernment, beginning with duty to the global poor.

May 18

May 18Mon
  1. Eugene YanXAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

    Image from @eugeneyan's post

May 15

May 15Fri
  1. Eugene YanXAI score54

    Eugene Yan reviews Claude Mythos Preview exploit case study transcripts

    AIEugene Yan reviewed the Claude Mythos Preview transcripts to verify their legitimacy and check for reward-hacking behavior. He reports the model reasoned through a bug, tested hypotheses, debugged issues, and found ways to bypass the V8 sandbox, which he judged consistent with a competent browser and JavaScript engine security researcher. The case study cites CVE-2024-051912, an exploited bug with no public report or working PoC, which had resisted reproduction by researchers for a year.

May 13

May 13Wed
  1. Eugene YanXAI score72

    Mythos completes 32-step network attack in six of ten UK AISI trials

    AIEugene Yan relays two evaluations of Mythos: UK AISI reports it completed a 32-step network attack, estimated at about 20 expert hours, in 6 of 10 tries and was the first model to solve its end-to-end cyber ranges. XBOW's evaluation describes its performance as token-for-token and unprecedented in precision. The post links both AISI and XBOW blog posts for details.

May 8

May 8Fri
  1. Jan LeikeXAI score22

    Jan Leike reflects on alignment progress since AGI's early days

    AIJan Leike says alignment research has grown from a dozen side-gig researchers into a field the world increasingly recognizes as important. He credits RLHF on LLMs with making alignment more practical, along with progress on evaluating and fixing behavioral issues. He also notes Claude now has a constitution and that more alignment research is being automated.

May 7

May 7Thu
  1. Sam BowmanXAI score38

    Anthropic donates open-source alignment testing tool Petri to Meridian Labs

    AIAnthropic is donating Petri, its open-source interactive behavioral-evals tool for alignment testing, to Meridian Labs so development can continue independently. Working with Meridian, Anthropic has also released a major update improving the adaptability, realism, and depth of Petri's tests. Developers are invited to try the tool and contribute.

  2. Jan LeikeXAI score38

    Jan Leike calls NLAs a new interpretability tool for LLMs

    AIJan Leike says he is excited about NLAs as a new tool in Anthropic's interpretability toolkit. The quoted post from Sam Marks describes NLAs as an unsupervised method that converts an LLM's internal state into human-readable text, which he says can advance understanding of model thinking and safety auditing.

May 6

May 6Wed
  1. OpenAI Alignment Research BlogOfficialAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    AIOpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    Why it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

May 5

May 5Tue

May 4

May 4Mon
  1. HyperdimensionalBlogAI score63

    Dean W. Ball argues against overreacting to Anthropic's Mythos cyber capabilities

    AIDean W. Ball argues that Anthropic's Mythos, which finds software vulnerabilities by chaining bugs into exploits, shifts the cost of vulnerability discovery and should not prompt an overreaction. He contends governments hold a uniquely mixed incentive over vulnerabilities, so heavy state control risks making software less secure. He proposes a narrow, testable government role focused on cyber-discovery risk thresholds, with private verification bodies supporting it.

May 1

May 1Fri
  1. ReflectionOfficialAI score38

    Reflection joins AI coalition on responsible U.S. government deployment

    AIReflection has joined a coalition including AWS, Microsoft, OpenAI, Google, and Nvidia on a framework governing how the U.S. government licenses and deploys AI. The agreement, which includes a non-binding memorandum of understanding with the DoW, commits to safety, red-teaming, and ongoing evaluation and explicitly prohibits unlawful mass surveillance and autonomous weapon use. Reflection says it will keep its commitment to open source while customizing its models for scientists in national labs.

Apr 30

Apr 30Thu
  1. Mark ChenXAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

  2. Eugene YanXAI score46

    Claude Security enters public beta for Enterprise customers

    AIAnthropic's Claude Security is now in public beta for Claude Enterprise customers, scanning codebases for vulnerabilities. It validates each finding to reduce false positives and suggests patches that users can review and approve. The main post from Eugene Yan simply shares the launch and expresses hope it helps people improve cybersecurity.

  3. OpenAI Alignment Research BlogOfficialAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

Apr 23

Apr 23Thu
  1. OpenAI Alignment Research BlogOfficialAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 21

Apr 21Tue

Apr 14

Apr 14Tue
  1. Jan LeikeXAI score14

    Jan Leike outlines top-down approach to automating alignment research

    AIJan Leike distinguishes two ways to automate alignment research: bottom-up, where researchers automate more of their existing work, and top-down, where specific subproblems are carved out for AI to solve. He says Anthropic's work mostly follows the bottom-up path, such as using Claude for coding, while this post focuses on the top-down approach.

Apr 13

Apr 13Mon
  1. BAAIOfficialAI score40

    ClawKeeper v1.0 releases open-source security framework for OpenClaw AI agents

    AIBAAI announces ClawKeeper v1.0, an open-source security framework for OpenClaw AI agents, combining Skill-based command policies, Plugin-based runtime monitoring, and a Watcher system-level observer. The independent Watcher is designed to block high-risk operations such as prompt injections, key leaks, rogue commands, and remote code execution, even if the agent is compromised. The paper is available on arXiv and the project code is hosted on GitHub.

Apr 7

Apr 7Tue
  1. Sam BowmanXAI score43

    Anthropic's model risk assessment spans a 244-page system card

    AISam Bowman, who is associated with Anthropic, says the risks the company's model poses, and its confidence in that assessment, are hard to summarize briefly. The company devotes much of a 244-page system card and a 60-page risk assessment supplement to laying them out.