Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs).
Key points
- Here we quantify how this LLM turnover affects the screening of scientific manuscripts.
- We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026.
- Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it.
- Vocabulary differences between versions largely track where detection transfers and where it fails.
Sources (1)
- [1]Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific WritingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:44 AM
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs).
Here we quantify how this LLM turnover affects the screening of scientific manuscripts.
Extractive summary: sentences quoted from the sources.