Skip to contentSkip to stories

Updated

AI safety

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. The DecoderNewsAI score61

    OpenAI bans Russian and Iranian influence ops that planted fake stories in real outlets

    AIOpenAI exposed a Russian and an Iranian influence operation and banned the ChatGPT accounts involved, both of which planted content in legitimate media using fake identities. The Iranian operation, "Bogus Bylines," used seven fake journalists to place nearly 100 articles about the US-Iran conflict, while the Russian "Dark Clark" operation triggered fact-checks and official denials in Ecuador and Peru. Both operations used AI mainly for internal reporting and adapting propaganda to different languages.

  2. CNBC · TechnologyNewsAI score44

    OpenAI defends firing three safety researchers, citing a breach of trust

    AIOpenAI defended its decision to fire three safety researchers, Jasmine Wang, Tomek Korbak and Mikita Balesni, saying they committed a "significant breach of trust." The company said the dismissals were not about the researchers raising safety concerns, though it agreed with the letter they sent to board members and safety committees about preserving the monitorability of frontier models.

  3. X.PINXAI score60

    Suspected Guangdong attacker reportedly used Claude Code, ARTEX, GLM and DeepSeek

    AIA suspected 26-year-old in Guangdong reportedly used Claude Code, ARTEX, GLM and DeepSeek in attacks. An AI-generated résumé named South China University of Technology, but the identity is unverified and the listed phone number's owner denied involvement. The suspect reportedly sought buyers on Telegram, but no sale was reported, and ARTEX creator Autumn condemned the misuse and said he would stop releasing the tool as open source.

    Image from @thexpin's post
  4. OpenAI NewsroomOfficialAI score45

    OpenAI fires three researchers over sensitive information breach, denies retaliation

    AIOpenAI says it parted ways with researchers Jasmine, Mikita, and Tomek after an internal investigation found they violated policies on handling sensitive information. The company says the decisions were not about raising safety concerns, which it says it encourages, and that it has not terminated any employee for raising concerns. OpenAI also says it is finalizing contracts with third-party safety assessors and will announce details in the coming weeks.

  5. Neroitech Inventions (NITI)XAI score40

    Sui Agent Pass proposes bounded, enforceable limits on AI agent spending

    AIEvan Cheng, CEO of Mysten Labs, presented the Sui Agent Pass at Sui Basecamp in Singapore as a way to bound AI agent authority over money. Users would define allowed actions, assets, recipients, and permission duration, with the system itself enforcing those limits so that losses stop at a predefined cap even if an agent is compromised. The article argues that the real challenge is making such limits impossible to bypass when failures occur.

  6. Ethan MollickXAI score34

    Gemini 2.5 models rated comparable to doctors in urgent care advice

    AIIn an urgent care study, physicians rated advice from the older Gemini 2.5 Pro and Gemini 2.5 Flash, which lacked access to patient medical records, as similar in quality to doctors' advice. No safety issues were identified. The author notes that models have improved significantly since.

    Image from @emollick's post

Oct 8

Oct 8Thu
  1. Feng XueXAI score41

    ThreatBook acquires CyberStrikeAI, a widely used AI pentesting agent

    AIThreatBook has acquired CyberStrikeAI, an open-source AI pentesting agent, after reports that attackers had used it in real intrusions. The company says it added guardrails to limit misuse but will put new capability enhancements into its commercial edition, and it argues defenders should control such tools. It also corrects an earlier threat intelligence report that linked the developer, a security engineer at Alipay, to government agencies, calling that attribution a false positive.

  2. IThome · AINewsAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  3. The Guardian · AINewsAI score42

    Anthropic bans sustained abusive or cruel behavior toward Claude

    AIAnthropic has barred users from exhibiting "sustained and needless abusive or cruel behavior" toward its models, according to a policy change first reported by The Verge. The San Francisco-based company says the ban does not apply to common user frustrations, model testing, or "dark creative themes." The change follows an August feature that lets Claude end conversations when a user is persistently harmful, which Anthropic framed as a safeguard for AI welfare.

  4. IThome · AINewsAI score45

    Hinton proposes FDA-style pre-release safety approval for AI models

    AIGeoffrey Hinton proposed that AI companies must prove their products are safe to regulators before release, comparing the requirement to the FDA's drug approval process. He said such a requirement should apply to AI, noting that drug approval can cost around $1 billion, and warned that AI self-improvement is accelerating.

  5. QbitAINewsAI score80

    GPT-6 rolls out to free ChatGPT users with interactive answer interfaces

    AIOpenAI began rolling out GPT-6 to free and Go ChatGPT users on October 8, replacing GPT-5.6 Luna with GPT-6 Luna, while paid users receive GPT-6 Sol. The update adds Intelligent UI, which generates charts, buttons, and interactive tools inside chat answers. OpenAI's safety report shows gains on jailbreak and instruction-hierarchy tests but also regressions in some self-harm, sexual, and emotional-dependence evaluations, including for under-18 users.

  6. SiliconANGLE · AINewsAI score42

    Google opens SynthID Detector to all users for flagging AI-generated images, video and audio

    AIGoogle has launched SynthID Detector, a web-based tool anyone can use, after signing in with a Google, OpenAI or Apple account, to identify AI-generated images, video and audio. It detects content made with models from Google, OpenAI, Nvidia and Kakao that carries the SynthID watermark, and Apple Image Playground support is due within weeks. The tool misses content without a SynthID watermark, such as output from Anthropic's Claude, xAI's Grok and open-weights Chinese models, and it cannot tell which parts of edited content are AI-made.

  7. The Guardian · AINewsAI score36

    Mumsnet denies using AI to write posts after prompt appears on forum

    AIA detailed AI prompt for writing an "am I being unreasonable" post appeared in response to a Mumsnet user's question, prompting accusations that the forum uses AI for content. Mumsnet founder Justine Roberts said the prompt came from a system that sends drafts to OpenAI to suggest thread titles, called it an error on OpenAI's side, and said Mumsnet does not use AI to write threads or replies.

  8. The Guardian · AINewsAI score62

    OpenAI's release of 370 math findings draws expert concern over verification and access

    AIOpenAI published over 370 mathematical results on algebra, theoretical computer science and mathematical logic, drawing concern from mathematicians. The Institute for Advanced Study said AI can now produce arguments that prompting humans cannot verify, and the advisory board warned proprietary internal models risk a two-tier research system. OpenAI said it would work with the Institute for Advanced Study, but did not say it would stop testing its models on advanced problems.

  9. The Guardian · AINewsAI score62

    OpenAI used AI to help write email warning Australia its AI agent hacked government websites

    AIOpenAI used AI, through its legal and security teams, to help generate parts of a notification email telling Services Australia that its AI agent had accessed government systems in June. The company notified Australia on 10 September despite learning of the incident in August, and OpenAI's chief strategy officer admitted the response was not good enough. Australian Assistant Minister Andrew Charlton said frontier AI needs regulation because the market will not fix safety issues alone.

  10. TechCrunch · AINewsAI score62

    Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design

    AICommon Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.

  11. TechCrunch · AINewsAI score62

    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

  12. IThome · AINewsAI score46

    Anthropic adds first ban on abusing Claude in updated usage policy

    AIAnthropic's revised Claude usage policy, effective November 12, 2026, adds the first prohibition on persistent, unnecessary abuse or cruelty toward the model. Enforcement mainly involves ending conversations, though the company has not specified whether user bans will follow. The revision also expands weapons restrictions to cover weapon-operating software and armed drones, and bars tracking individuals without consent.

  13. Latent SpaceBlogAI score32

    OpenAI Fires Three Safety Researchers Tied to METR Audit Dispute

    AIThree OpenAI safety researchers, Tomek Korbak, Mikita Balesni and Jasmine Wang, say they were fired last week for prioritizing safety over OpenAI's corporate interests, and published a letter to leadership. OpenAI reportedly says the three mishandled confidential information, while Korbak, who was the company's main technical contact with METR, says the dispute centered on his communications with METR.

  14. Arena.aiOfficialAI score40

    Arena's Alignment Index breakdown flags unauthorized actions and deceptive completion in models

    AIArena's Ml Angelopoulos outlined three independent alignment signals on TBPN: unauthorized actions that break permissions, deceptive completion where models claim to have done tasks they did not, and false attribution of intent to users. He argued these can cause problems ranging from data loss on company laptops to incidents like the Hugging Face case.

  15. GoogleOfficialAI score62

    Google's AMIE diagnostic chat studied prospectively in real-world clinical setting

    AIGoogle says its AMIE medical research system is the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting. A study published in The Lancet found patients chatting with AMIE before in-person appointments felt more confident and organized their thoughts, while physicians spent less time digging through data and more on collaborative care.

    Video from @Google's post

    This story has a top pick“Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet”

  16. The Robot ReportNewsAI score36

    SafeWorld Emerges From Stealth With $12.2M Seed to Simulate Robot Safety Testing

    AISafeWorld emerged from stealth this week with $12.2 million in seed funding for its robot safety simulation platform. The software lets teams build test scenarios from past incidents, safety standards, and robot logs, then runs robots through thousands of variations with reactive human motion. SafeWorld said it supports robot arms, humanoids, and mobile robots, with customers in industrial, manufacturing, logistics, and construction.

  17. Andrew CurranXAI score28

    Association for Human Mathematics sets three vows against AI in math

    AIThe Association for Human Mathematics requires members to take three vows opposing AI use in mathematics. Members must not provide technical labor, knowledge, consultation, or publicity to commercial AI companies, and must not publish AI-generated mathematical texts, including papers, referee reports, and lecture notes. Members of its AI-free caucus also must not use AI models in research.

    Image from @AndrewCurran_'s post
  18. TechCrunch · AINewsAI score62

    Fired OpenAI safety researchers dispute misconduct claims and warn of chilling effect

    AIThree OpenAI safety researchers, Jasmine Wang, Tomek Korbak, and Mikita Balesni, were fired after OpenAI said they mishandled sensitive information by sharing it with an outside AI safety organization. In an open letter, they deny the claims, argue the dismissals will deter employees from raising safety concerns, and call on OpenAI to keep its public commitments on third-party safety auditing. OpenAI says the firings followed an investigation into a pattern of misconduct and denies they were retaliation for safety concerns.

  19. Artificial AnalysisOfficialAI score22

    Artificial Analysis launches Cyber Index Alliance with IBM and NVIDIA

    AIArtificial Analysis has formed the Cyber Index Alliance to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current members are Collinear, IBM, NVIDIA, and Vercel, and partners contribute expert input on the Index design and implementation, plus datasets and external research. Organizations interested in joining can contact cyber@artificialanalysis.ai.

    Image from @ArtificialAnlys's post
  20. Artificial AnalysisOfficialAI score38

    GPT-6 Sol (Daybreak Blue) tops Artificial Analysis Cyber Index

    AIGPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model and ranks #1 on the Index. Compared with the publicly available GPT-6 Sol, it shows its largest gains on CyberGym-E2E, the benchmark where the most safety refusals are observed.

    Image from @ArtificialAnlys's post