The Scientific Literature is Poisonous to LLMs
Seriously, what would you expect?
We can think of no corpus of text more toxic to the training of useful LLMs than the 21st century scientific literature. Imagining a two-by-two matrix with the axes honest vs. dishonest and right vs. wrong, the scientific literature is splashed haphazardly across all four boxes. Papers often belong in multiple quadrants at once, and not always because of the contributions of different authors. Even worse, papers in any quadrant might deliberately masquerade as work in any other quadrant. Consider, for example, the ordinary scientist who must dress up an honest experiment in a paper full of superlatives and explain that experiment with an incorrect but popular theory so that she can hope to compete for a journal spot against the fraudster down the hall who is free to make up whatever data is most convenient. The result is a minefield of half-truths, convenient omissions, and outright lies that can’t be distinguished from the innocent facts and findings around them. And we won’t even talk about the writing style.
There’s empirical evidence for the negative effect of the academic literature on LLM performance. In 2024 a research team spanning MIT, Cornell, Carnegie Mellon, Google, and OpenAI examined the effects of removing different text corpora from the training data on the performance of an LLM after training (holding the LLM’s structure constant)1. They found that removing ArXiv, PhilPapers, and NIH ExPorter from the training corpus improved the LLM’s performance at answering academic questions and its average performance across all benchmarks while also making the model less likely to generate toxic output. The study is not conclusive on this point because removing Pubmed reduced model performance somewhat, but the mere fact that LLM performance could be improved by wholesale removal of large sets of scientific articles (pre-slop no less!) is pretty suggestive. We wonder whether this result holds up in 2026, but we’d bet that it does.
Scientific metadata is no help either. Once authorship and citations started being used as a metric, all the signal was gamed out of them. Furthermore, high profile institutions and researchers are not immune to misconduct, nor are any particular fields.
This has important practical consequences for using AI to do science. Human scientists use their informal networks to communicate reputational information, verify work through visits and replication, and generally maintain idiosyncratic personal pictures of what work is reliable and what work matters. Because AI agents do not participate in social relationships they cannot access this social metadata except through humans.
This leaves AI for Science efforts with a strange set of choices. They can either create AI agents that human scientists socialize with, frequently record human scientists in private settings, or create an incentive for scientists to provide a steady stream of gossip. Perhaps one way to bootstrap this effort would be to create AI assistants that allow for more scientific work to be done in informal settings. This would amount to surveillance, but perhaps could be worth it for a big enough improvement in our quality of life.
We write Reinvent Science to broaden the conversation around science and science funding, and we rely on you to help us reach as many readers as possible. Please support our work by subscribing and sharing.
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito. 2024. “A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.” In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human


