Skip to content
View original post on X: Rohan PaulX· 46/100AI score46/100

AI-text detectors miss most rewrites from newer model generations

AISummary

A Tokyo Metropolitan University paper finds AI-text detectors trained on a vendor's older models caught over 99% of rewrites before a generation change but only 3.8% after it.

Pangram missed 79.8% of abstracts rewritten by Meta's Muse-Glimmer while catching 93.5% of GPT-5 rewrites, and flagged just 1 of 5,000 human abstracts. The paper's authors say anyone relying on such detectors should re-test them with each model release.

Post on XView on X
Rohan PaulVerified on X
@rohanpaul_ai

This paper finds that AI-text detectors trained on a vendor's older models caught over 99% of rewrites before a generation change and only 3.8% after it.

Detectors that screen scientific papers for AI writing can stop working when a new LLM generation arrives, so anyone relying on them should re-test them with every model release.

Rohan Paul@rohanpaul_ai
Pangram, the AI-text detector, missed 79.8% of scientific abstracts rewritten by Meta's Muse-Glimmer, while flagging just 1 of 5,000 human abstracts. In a new paper from Tokyo Metropolitan University, reseaerchers find the share of AI-rewritten abstracts that Pangram misses depends strongly on the LLM version Shows that it caught 93.5% of GPT-5 rewrites but missed 79.8% from another new model. Its miss rate depended mostly on which model did the rewriting.
View quoted post on X

Source: Rohan Paul · x.comPublished · added here