Years of iterating against the same benchmarks should, by textbook logic, produce overfitting. It largely doesn't.
New research explains why: strategies that generalize can be expressed in too compact a form to allow memorization, while the ones that overfit don't survive a compression bottleneck. https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit?utm_campaign=why-dont-machine-learning-research-agents-overfit&utm_medium=organic-asw&utm_source=twitter&utm_content=2026-09-10-why-dont-machine-learning-research-agents-overfit&utm_term=2026-september
