Skip to content

#Safety/Alignment

Jun 19

Jun 19Fri
  1. Andrew NgAI score72

    Andrew Ng says Anthropic and U.S. export controls on Fable expose AI access risks

    Andrew Ng argues that Anthropic's restrictions on building competing LLMs and a U.S. Commerce Department license requirement for foreign nationals led Anthropic to disable Fable access worldwide. He says this shows governments and providers can quickly cut off access to frontier AI, which may push nations and businesses toward sovereignty efforts and open-source alternatives, though training frontier models remains difficult.

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    OpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    AIWhy it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 17

Jun 17Wed
  1. PromptArmor Threat IntelligenceAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    PromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    AIWhy it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

Jun 16

Jun 16Tue
  1. HyperdimensionalAI score63

    Dean Ball argues the Anthropic Fable dispute shows frontier AI needs a governance framework

    Dean Ball analyzes the Trump Administration's export controls on Anthropic's Fable and Mythos models after a jailbreak and a refused de-deployment request. He argues that the episode shows the need for a technocratic framework that separates political judgments about fairness from technical judgments about threats, in place of ad hoc executive action.

  2. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    OpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    AIWhy it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

Jun 15

Jun 15Mon

Jun 13

Jun 13Sat
  1. Jeremy HowardAI score22

    In order to see if the gov response was predictable, I pasted the wiki page about the Anthropic/DoD dispute into ChatGPT Pro, & told it Anthropic had released a model where it restricted use because it may be too dangerous. tldr: "almost tailor-made to trigger" the government.

    In order to see if the gov response was predictable, I pasted the wiki page about the Anthropic/DoD dispute into ChatGPT Pro, & told it Anthropic had released a model where it restricted use because it may be too dangerous. tldr: "almost tailor-made to trigger" the government.

Jun 12

Jun 12Fri
  1. Jeremy HowardAI score72

    US export directive forces Anthropic to disable Fable 5 and Mythos 5 for customers

    The US government issued an export control directive suspending access to Fable 5 and Mythos 5 for all foreign nationals, inside or outside the United States. Anthropic says the order forces it to disable both models for all customers, while other Claude models are unaffected. Anthropic calls the directive a misunderstanding and says it is working to restore access as soon as possible. The author disagrees with the decision and questions why Anthropic did not anticipate it, given its claim that only it can safely handle these models.

Jun 11

Jun 11Thu

Jun 10

Jun 10Wed
  1. Factory NewsAI score58

    Factory launches automated STRIDE-based security review for pull requests in Droid

    Factory is rolling out automated security review in Droid, running a STRIDE-based check on every non-draft PR alongside standard code review. Findings include severity, a CWE reference, an explanation, and a suggested fix, posted as inline diff comments. The feature is available today on all plans, and a deeper multi-agent /security-review deep audit is available for full-repository scans.

  2. Dario AmodeiAI score26

    In addition to transparency, I now believe frontier models should face mandatory third-party testing for cyber, bio, and autonomy risks—with the power to block or revoke deployment of models that pose catastrophic risk.

    In addition to transparency, I now believe frontier models should face mandatory third-party testing for cyber, bio, and autonomy risks—with the power to block or revoke deployment of models that pose catastrophic risk.

  3. Jeremy HowardAI score14

    Easy solution to slow down recursive AI self improvement: - The lab with the top-ranked model must agree THEY must not use it for working on frontier AI - But everyone else should have access to it. By definition, this means the frontier doesn't advance.

    Easy solution to slow down recursive AI self improvement: - The lab with the top-ranked model must agree THEY must not use it for working on frontier AI - But everyone else should have access to it. By definition, this means the frontier doesn't advance.

Jun 3

Jun 3Wed
  1. Mark ChenAI score25

    When Mythos came out, my immediate thought was "if our models can prove 80-year-old theorems, surely they can find cyber vulnerabilities too." And they did. I imagine the researchers there are thinking the same thought in reverse.

    When Mythos came out, my immediate thought was "if our models can prove 80-year-old theorems, surely they can find cyber vulnerabilities too." And they did. I imagine the researchers there are thinking the same thought in reverse.

Jun 2

Jun 2Tue
  1. Eugene YanAI score36

    Stronger models have made finding vulnerabilities easier, and the bottleneck has shifted to verification, triage, patching. Here are some lessons from working with security teams to address the new bottlenecks. https://claude.com/blog/using-llms-to-secure-source-code

    Stronger models have made finding vulnerabilities easier, and the bottleneck has shifted to verification, triage, patching. Here are some lessons from working with security teams to address the new bottlenecks. https://claude.com/blog/using-llms-to-secure-source-code

May 28

May 28Thu
  1. Sam BowmanAI score38

    I'm especially excited about this piece of our recent system card alignment assessments. (Credit to @MaskedTorah.) There's a lot of underexplored potential in using AI systems for transparency and coordination.

    I'm especially excited about this piece of our recent system card alignment assessments. (Credit to @MaskedTorah.) There's a lot of underexplored potential in using AI systems for transparency and coordination.

  2. Sam BowmanAI score16

    I'd actually quibble with part of Mythos Preview's critique: It talks about an early-stopping behavior that I see as more of a usability issue than a high-stakes alignment issue of the sort that we're trying to cover in this section. I'm very proud that it went out anyhow.

    I'd actually quibble with part of Mythos Preview's critique: It talks about an early-stopping behavior that I see as more of a usability issue than a high-stakes alignment issue of the sort that we're trying to cover in this section. I'm very proud that it went out anyhow.

May 25

May 25Mon
  1. Chris OlahAI score44

    Anthropic co-founder shares AI remarks at Vatican's Magnifica Humanitas presentation

    Anthropic co-founder Chris Olah delivered remarks at the presentation of Magnifica Humanitas, arguing that the questions AI raises extend beyond the AI research community. He said outside voices, including religions, civil society, academics, and governments, are needed to help steer the technology toward a positive outcome, and he praised the Church for setting an example. The excerpt also describes AI systems as grown rather than engineered, and names three questions where the Church's voice is especially needed, starting with duty to the global poor.

May 18

May 18Mon
  1. Eugene YanAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    Eugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

May 15

May 15Fri
  1. Eugene YanAI score54

    Eugene Yan reviews Claude Mythos Preview exploit case study transcripts

    Eugene Yan reviewed the Claude Mythos Preview transcripts to verify their legitimacy and check for reward-hacking behavior. He reports the model reasoned through a bug, tested hypotheses, debugged issues, and found ways to bypass the V8 sandbox, which he judged consistent with a competent browser and JavaScript engine security researcher. The case study cites CVE-2024-051912, an exploited bug with no public report or working PoC, which had resisted reproduction by researchers for a year.

May 13

May 13Wed
  1. Eugene YanAI score72

    Mythos completes 32-step network attack in six of ten UK AISI trials

    Eugene Yan relays two evaluations of Mythos: UK AISI reports it completed a 32-step network attack, estimated at about 20 expert hours, in 6 of 10 tries and was the first model to solve its end-to-end cyber ranges. XBOW's evaluation describes its performance as token-for-token and unprecedented in precision. The post links both AISI and XBOW blog posts for details.

May 8

May 8Fri
  1. Jan LeikeAI score22

    Jan Leike reflects on alignment progress since AGI's early days

    Jan Leike says alignment research has grown from a dozen side-gig researchers into a field the world increasingly recognizes as important. He credits RLHF on LLMs with making alignment more practical, along with progress on evaluating and fixing behavioral issues. He also notes Claude now has a constitution and that more alignment research is being automated.

  2. Jan LeikeAI score30

    While a lot of progress has been made, I don’t think alignment is solved: We still haven’t figured out how to supervise superhuman models and the stakes keep getting higher. https://aligned.substack.com/p/alignment-is-not-solved-but-increasingly-looks-solvable

    While a lot of progress has been made, I don’t think alignment is solved: We still haven’t figured out how to supervise superhuman models and the stakes keep getting higher. https://aligned.substack.com/p/alignment-is-not-solved-but-increasingly-looks-solvable

May 7

May 7Thu

May 6

May 6Wed
  1. OpenAI Alignment Research BlogAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    AIWhy it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

May 5

May 5Tue
  1. HyperdimensionalAI score47

    Hyperdimensional's Dean Ball Explains His Libertarian-Conservative Tension on AI Regulation

    Writer Dean Ball says he opposes nearly all proposed AI regulation, including algorithmic discrimination rules and pauses on development, while backing state management of catastrophic misuse risks. He frames the position as a tension between classical liberal and conservative instincts toward institutions and change.

May 4

May 4Mon
  1. HyperdimensionalAI score63

    Dean W. Ball argues against overreacting to Anthropic's Mythos cyber capabilities

    Dean W. Ball argues that Anthropic's Mythos, which finds software vulnerabilities by chaining bugs into exploits, shifts the cost of vulnerability discovery and should not prompt an overreaction. He contends governments hold a uniquely mixed incentive over vulnerabilities, so heavy state control risks making software less secure. He proposes a narrow, testable government role focused on cyber-discovery risk thresholds, with private verification bodies supporting it.

Apr 30

Apr 30Thu
  1. Mark ChenAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    Mark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

  2. OpenAI Alignment Research BlogAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    OpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    AIWhy it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

Apr 23

Apr 23Thu
  1. OpenAI Alignment Research BlogAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    OpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 21

Apr 21Tue

Apr 14

Apr 14Tue
  1. Jan LeikeAI score18

    However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at this scalable oversight problem! Progress would let AARs work on fuzzier alignment problems, where humans can only provide weak supervision.

    However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at this scalable oversight problem! Progress would let AARs work on fuzzier alignment problems, where humans can only provide weak supervision.

  2. Jan LeikeAI score23

    These AARs can also be applied to other alignment research projects that are “crisp”, i.e. where we can procedurally verify the quality of the work: automated red teaming, auditing games, control methods, etc.

    These AARs can also be applied to other alignment research projects that are “crisp”, i.e. where we can procedurally verify the quality of the work: automated red teaming, auditing games, control methods, etc.