PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End
Original title (Chinese)
ChatGPT踢到铁板了!能破解千禧数学难题,但论文复现率低至13.98%?
AISummary
UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.
Source: QbitAI · qbitai.comPublished · added here