This paper finds that AI-text detectors trained on a vendor's older models caught over 99% of rewrites before a generation change and only 3.8% after it.
Detectors that screen scientific papers for AI writing can stop working when a new LLM generation arrives, so anyone relying on them should re-test them with every model release.
