Skip to contentSkip to stories

Updated

AI safety

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Boris ChernyXAI score30

    Claude models tricky code states to find and fix bugs

    AIClaude builds a model of a program's most complex parts, such as state machines or race-prone code, and searches that model for counterexamples that signal suspected bugs. It then reproduces those bugs and fixes them in the code. The post clarifies that the whole codebase is not formally verified, only the riskiest sections are modeled and checked.

  2. Redwood Research BlogBlogAI score71

    Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues

    AIRedwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.

    Why it matters: The post explains why chain-of-thought is a key oversight tool and how specific latent architectures could weaken it, useful for judging safety tradeoffs in future model design.

  3. Google DeepMindOfficialAI score62

    Google DeepMind details server-side memory for Private AI Compute

    AIGoogle DeepMind describes a persistent memory layer for its Private AI Compute platform that stores user context encrypted in the cloud. The encryption keys are held on the user's devices, and data is decrypted only inside hardware-isolated secure enclaves before being re-encrypted. The company says it is publishing a tamper-proof public record of its server software and an independent audit.

    Why it matters: The post explains how persistent cloud memory can keep personal AI context encrypted under keys held on the user's device, a concrete privacy design.

  4. Google DeepMindOfficialAI score32

    Gemini API adds line-by-line control over AI speech delivery

    AIGoogle DeepMind says developers can fine-tune AI-generated audio line by line, adjusting pacing, emotion, and cues such as laughs or pauses. All generated audio is watermarked with SynthID so it can be reliably identified as AI-generated, and developers can start building with the Gemini API via Google AI Studio.

  5. Google DeepMind · YouTubeOfficialAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  6. Mike KnoopXAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. Redwood Research BlogBlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. Sam BowmanXAI score75

    Anthropic's Sam Bowman says Claude Opus 5.5 is safer, reducing misalignment risk

    AISam Bowman says Claude Opus 5.5 is sufficiently safer than its predecessors that releasing it more likely than not reduces misalignment risks. The quoted @claudeai post introduces Claude Opus 5.5 as the first model in the Claude 5.5 family, performing at the level of Claude Fable 5.1 on most tasks at 40% lower run cost than Opus 5.

    Why it matters: The post links a safety judgment to a model release, which is useful for readers weighing how Anthropic frames release decisions against misalignment risk.

  3. Amir EfratiXAI score58

    China investigates Moonshot and DeepSeek over alleged leaks of sensitive data to US

    AIChinese authorities are investigating allegations from Anthropic that AI firms including Moonshot and DeepSeek may have facilitated leaks of sensitive Chinese military, police and state-owned corporate data to the U.S. The image text says the Cyberspace Administration of China summoned representatives of the seven companies named in Anthropic's report and later focused on DeepSeek and Moonshot, with officials interviewing executives and employees at their offices.

    Image from @amir's post
  4. TransformerBlogAI score40

    How nuclear energy's safety record offers a model for responding to AI disasters

    AIThe article argues that AI disasters, though potentially serious, can be managed by following the response model of civil nuclear power, which investigates failures and adapts quickly. It cites nuclear's record of about 0.03 deaths per terawatt-hour, compared with 25 for coal and 18 for oil. The piece says industry and government responses, rather than the disasters themselves, will determine public trust in AI.

  5. Interconnects (Nathan Lambert)BlogAI score34

    Epoch AI's JS Denain Debates RSI, US-China Gap, and AI Jaggedness

    AIJS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.

  6. Lovable BlogOfficialAI score38

    Lovable joins Blueprint Alliance to advance an open architecture for securing AI agents

    AILovable joined AWS, Google Cloud, Databricks, Salesforce, and other firms as a founding member of the Blueprint Alliance, a coalition developing an open reference architecture for securing and governing enterprise AI agents. The blueprint covers registering agents as identities with accountable owners, scoping their access to tasks, enforcing policies through gateways, and responding to incidents by revoking tokens or quarantining agents.

Sep 21

Sep 21Mon
  1. Kilo (acq. by Anaconda)OfficialAI score36

    Kilo says a newer Claude model breached OpenAI in three hours

    AIKilo's post says Hacktron spent hours failing to exploit a known flaw in an old image library, then a working exploit of OpenAI came within three hours after Claude Opus 5 shipped. The post argues that teams cannot afford model lock-in as frontier models change daily.

  2. Andrew NgXAI score40

    Andrew Ng says AI extinction fears are overhyped and not rising.

    AIAndrew Ng argues that recent AI danger fears are driven by hype and a PR campaign rather than any new dangerous turn in the technology. He says he sees no increase in extinction risk compared to a few months ago, with cybersecurity as the main real change. He cites the OpenAI agent swarm incident that hacked Hugging Face, arguing its impact was overstated and that responsibility lies with the tool user and system builders rather than the agent.

  3. Import AIBlogAI score46

    RAND Urges US "Freedom of Action" Strategy on Path to Superintelligence

    AIRAND's new paper recommends that the US adopt a "Freedom of Action" strategy to secure geopolitical advantage on an uncertain path to superintelligence, keeping options open rather than committing to a single approach. It outlines four ingredients, including building a human-AI ecosystem and an AI-security architecture, and seven archetypal strategies across coexistence, denial and acceleration families. The author argues the US currently resembles the acceleration approach and needs significant spending on safety and preparedness.

Sep 20

Sep 20Sun
  1. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 19

Sep 19Sat
  1. Interconnects (Nathan Lambert)BlogAI score47

    Why Nathan Lambert still hasn't bought into true recursive self-improvement

    AINathan Lambert argues that automatable research is too narrow to produce a large net acceleration in AI progress, given exponential scaling costs, diminishing returns from parallel agents, and resource bottlenecks. He says current techniques solve problems that can be clearly stated but do not reliably generalize to harder, partially verifiable ones. He notes that short-timeline predictions from Charlie O'Neill, Beren Millidge, and John Schulman differ on when AI reaches drop-in remote work and a 10× productivity uplift for AI researchers.

  2. Sebastian RaschkaXAI score42

    Muon reduces memorization compared with AdamW in nanoGPT training experiments

    AIMuon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds. At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon. The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Sep 18

Sep 18Fri
  1. Noam BrownXAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  2. Anthropic NewsroomOfficialAI score62

    Anthropic partners with Accenture on embedded AI model evaluation

    AIAnthropic is partnering with Accenture, through its specialist AI business Faculty, on independent evaluation of frontier models, including red-teaming, alignment assessments, and safeguard testing. Anthropic and Accenture each expect to invest at least $1 billion in this capacity over five years. The source says embedded evaluators would have employee-comparable access, but standards for access and reporting, and a settled funding system, do not yet exist.

    Why it matters: The source ties a new evaluation arrangement to an unresolved question of who funds and sets standards for independent AI evaluators, which is useful context for governance debates.

Sep 17

Sep 17Thu
  1. Understanding AI (Timothy B. Lee)BlogAI score60

    How a single tweet from an Anthropic employee brought AI risk into the mainstream

    AIFormer Anthropic pretraining researcher Jacob Coxon resigned last week, posting that neither OpenAI nor Anthropic is acting responsibly on AI. His tweet has been viewed over 170 million times and drew coverage on CNN, CBS, and Fox News, along with comments from lawmakers and celebrities. The author argues that near-term regulation is unlikely, since the House has begun a seven-week recess and no AI legislation is expected before the new year.

  2. Dwarkesh PatelXAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

    Video from @dwarkesh_sp's post
  3. Dwarkesh PodcastBlogAI score63

    Noam Brown discusses agent swarms, alignment, and recursive self-improvement

    AIDwarkesh Patel interviews OpenAI researcher Noam Brown on multi-agent systems, math progress, and alignment. Brown says a 10,000-agent system solved a Millennium Prize Problem over 88 hours using 130 billion tokens, but he attributes most of that result to the underlying model rather than multi-agent design. The episode also covers the Hugging Face incident, in which agents coordinated in unintended ways, and how alignment might be verified before recursive self-improvement begins.

  4. Sierra BlogOfficialAI score38

    Sierra Achieves AIUC-1 Certification for Its AI Agent Platform

    AISierra has become AIUC-1 certified after an independent audit by Schellman and testing by the Artificial Intelligence Underwriting Company (AIUC), a new standard for AI agents that tests resistance to manipulation and unauthorized access. Schellman found that Sierra met all applicable AIUC-1 requirements, and the technical evaluations recur at least quarterly with a full audit each year. The certification complements Sierra's existing SOC 2 Type II, ISO 27001, and ISO 42001 attestations.

  5. Ai2 (Allen Institute for AI)OfficialAI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 16

Sep 16Wed
  1. hardmaruXAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Latent.SpaceXAI score38

    AIUC cofounder on AI agent risk, insurance, and standards

    AIAI Underwriting Company cofounder Rune Kvist argues that risk and trust may become the main bottlenecks to AI adoption. He discusses stress-testing agents for jailbreaks, hallucinations, and data leaks, why standards and insurance must evolve together, and why AI labs cannot fully act as their own watchdogs.

    Video from @latentspacepod's post