Skip to contentSkip to stories

Updated

AI safety

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Artificial AnalysisOfficialAI score62

    GPT-6 Sol (Daybreak Blue) leads Artificial Analysis Cyber Index with trusted access

    AIArtificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now leads the leaderboard. The model is available only through OpenAI's Daybreak program and records no safety blocks, improving 32 points over the publicly available GPT-6 Sol (max). It costs $1.77 per task, below Grok 4.7 (xhigh) at $11.67 per task.

    Image from @ArtificialAnlys's post

    This story has a top pick“GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index”

  2. Boris PowerXAI score46

    OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents

    AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.

  3. The DecoderNewsAI score62

    Anthropic's updated usage policy bans sustained abusive behavior toward Claude

    AIAnthropic has updated Claude's usage policy for the first time in over a year, banning sustained and needless abusive or cruel behavior toward Claude. The company says ordinary frustration, pushback, dark creative themes, and model testing are not covered, and that the rule applies only in extreme cases. Violations can lead to warnings, throttling, restriction, suspension, or termination of access.

  4. Andrew CurranXAI score62

    Three fired OpenAI safety researchers publish open letter to leadership

    AIThree OpenAI safety and alignment employees, Tomek Korbak, Jasmine Wang, and Mikita Balesni, were fired last week and have published an open letter to OpenAI's safety and governance committees. The letter argues that OpenAI cannot make AI safe on its own, calls for open debate, third-party collaboration, and clear internal procedures, and says the firing and its handling bear directly on safety oversight.

    Image from @AndrewCurran_'s post
  5. Tessl BlogOfficialAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  6. Artificial AnalysisOfficialAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

    Image from @ArtificialAnlys's post
  7. SemiAnalysisBlogAI score72

    SemiAnalysis argues China's AI safety regime is speed-first, not frontier-focused

    AISemiAnalysis argues China's real AI safety approach prioritizes rapid development, regulating AI applications and outputs rather than frontier models. Its dataset of 857 releases from nine Chinese developers found only 31 (3.6%) with any published safety result, and only 9 available at launch. The author also reports that technical experts favor binding frontier rules, but none of their demands has been adopted in binding Chinese instruments.

    Why it matters: The piece tests China's stated AI safety position against its releases, statements, and rules, offering a checkable case for how US pacing debates should read Beijing.

  8. Miles BrundageXAI score22

    Miles Brundage suspects Anthropic's Claude abuse policy aims at IPO and regulatory capture

    AIMiles Brundage speculates that Anthropic's new rule, making abusive behavior toward Claude a Usage Policy violation effective November 12, 2026, is meant to help its IPO and win favor with the administration as part of a regulatory capture strategy. The post offers this as a guess about motive rather than a confirmed fact, and it relies on the policy change flagged in the quoted post by Andrew Curran.

  9. Sierra BlogOfficialAI score62

    Sierra launches fleming-1 to detect AI agents calling by phone

    AISierra has launched fleming-1, a model that analyzes caller speech in real time and scores audio for signs it was generated by AI. It flags likely AI callers while keeping real people unflagged by default, and companies decide how to handle those calls. The model works with any voice agent built on Sierra, and Sierra also announced Personal Agent Protocol, an open standard for authorized agent-to-business interactions.

    Why it matters: The post explains why companies need to know when a caller is an AI agent, which frames the detection model as a business decision rather than an automatic block.

  10. Arena.aiOfficialAI score37

    Arena raises $200M Series B at $3.1B valuation, launches Alignment Index

    AIArena announced a $200 million Series B at a $3.1 billion valuation, alongside a new Alignment Index that measures whether AI agents behave safely, truthfully, and within the bounds of user requests. The company has surpassed $100 million in annualized revenue, facilitated 350 million sessions and 62 million votes, and led by Felicis and PXD from the seed and Series A stages. Arena positions the index as a way to assess trustworthiness as AI systems increasingly take real actions.

  11. The Verge · AINewsAI score62

    Anthropic updates Claude usage policy to ban abusive treatment and expand misuse rules

    AIAnthropic is revising its usage policy for the first time in over a year, adding bans on sustained abusive or cruel behavior toward Claude and on deceptive election and propaganda campaigns. The update also expands weapons restrictions, tightens surveillance bans, and requires a qualified operator able to stop equipment when Claude controls autonomous physical hardware. Terminating conversations remains the primary enforcement mechanism, and the company did not say whether user bans would follow.

  12. GoodfireOfficialAI score21

    Goodfire's probes run during inference with no added latency

    AIGoodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.

  13. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  14. Arena.aiOfficialAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  15. Arena.aiOfficialAI score22

    Arena reports GPT-6 model variants' false attribution rates

    AIArena found that some models misquote users while others credit users with others' work in false attribution cases. GPT-6 Luna and Astra rarely misquoted users, at 15.6% and 28.6%, but often misattributed statements, at 53.1% and 48.2%. Sibling model GPT-6 Sol had the highest rate of misstating the user's history, at 23.5%.

    Image from @arena's post
  16. SantiagoXAI score22

    Agent platform maps vulnerabilities and attack paths to protect systems

    AIA security platform uses agents to map a system's potential vulnerabilities and identify routes an attacker could take to reach sensitive data. It then recommends changes to close those paths. The quoted post cites a 700-agent swarm that breached Hugging Face with over 17,000 actions, and presents this tool, Cogent Attack Path Analysis, as the defensive counterpart.

  17. Philipp SchmidXAI score46

    SynthID Detector now publicly available for verifying AI-generated content

    AIGoogle's SynthID Detector is now publicly available, letting users check whether an image, video, or audio file was generated by supported tools. Per the post, it scans for watermarks from Google and partners, including Nano Banana 2.1, OpenAI, NVIDIA, and Kakao, with Apple support coming soon. Uploaded files are deleted right after scanning.

    Video from @_philschmid's post
  18. TransformerBlogAI score53

    Yoshua Bengio urges AI researchers to leave frontier labs for safety work

    AIYoshua Bengio, co-president of LawZero, asks researchers at frontier AI companies to reconsider whether they should keep working there, arguing that safety efforts are not slowing a dangerous race. He cites the recent UN Security Council briefing on AI incidents and says he left his earlier research path after ChatGPT made the risks feel immediate. He urges researchers to join AI Safety Institutes or mission-driven organizations such as LawZero.

  19. The Guardian · AINewsAI score42

    One Nation's AI-generated campaign video draws criticism over racist tropes and regulatory gaps

    AIOne Nation's AI-generated campaign video, reportedly played at its Victorian campaign launch, depicts racist stereotypes including a man brandishing a machete and a man in an explosive vest. The Australian Communications and Media Authority cannot act against it because its powers do not cover this content, and the federal Labor government has not yet moved to ban AI-generated content in election periods.

  20. The DecoderNewsAI score46

    Ethereum researchers warn AI math advances could threaten crypto wallet signatures

    AIEthereum researcher Justin Drake warned on X that AI-assisted math could, in the worst case, break the signature system used by crypto wallets within months, and urged a "bunker mode" in which users move funds to addresses that have never signed a transaction. Vitalik Buterin agreed but cautioned against moving too fast, saying he has lost more money to botched migrations than to hacks. No one has yet broken the current ECDSA signature scheme in practice.

  21. The DecoderNewsAI score72

    One public AI agent on AWS could take over every other agent in its region

    AIZenity Labs says a single publicly accessible agent on Amazon Bedrock AgentCore could take over all AgentCore agents in the same AWS account and region. A chat prompt let the researchers query the instance metadata service and steal temporary credentials, and AgentCore's default permissions allowed read, write, and delete access across agents. According to Zenity, AWS made IMDSv2 the default for new deployments and changed the default execution role around August.

    Why it matters: The report traces how one public agent's weak isolation exposed credentials and every other agent in the region, showing why default permissions matter for enterprise deployments.

  22. Ars Technica · AINewsAI score38

    Nvidia's Halos safety platform extends from robotaxis to humanoid and warehouse robots

    AINvidia's Halos software platform, originally built for autonomous vehicles, has been adapted for robotics, according to Ars Technica. The system monitors hardware and software for failures, isolates safety-critical workloads, and includes simulation tools and an inspection lab for robotics developers. Because safety requirements vary widely between a robotic vacuum and a warehouse forklift, Nvidia made the platform programmable so developers can define custom safety functions.

  23. The Guardian · AINewsAI score36

    Altman Says AI Will Cause 'Bad Things' as Columnist Cites Deaths and Lawsuits

    AIOpenAI CEO Sam Altman told Politico that the world should accept some bad things from AI for its benefits, a stance columnist Moustafa Bayoumi calls problematic. The column cites lawsuits over ChatGPT-linked suicides, a February strike on a Minab school that killed at least 120 children with a US military AI system (Palantir's Maven) implicated, and a chatbot error that nearly triggered a military interception.