> We wonder whether this result holds up in 2026, but we’d bet that it does.
I would be a little more impressed if the result were on something a little larger than a 0.02b-parameter T5 without even any error bars for tiny effects from removing a few % of the total training data. I hesitate to generalize from that to any contemporary LLMs like a Fable at ~10,000b-parameters.
The frontier labs routinely investigate data mixes in extreme detail, and I am unaware of any LLM which doesn't seem to have practically memorized Arxiv, or any reports from, say, the Chinese labs that they have found their own data mix ablations to support the claim that "The Scientific Literature is Poisonous to LLMs".
Dan here, one of the authors of this. Here's an attempt to go deeper. We'd be glad to receive feedback from you or other readers.
As far as we know, it is uncontroversial among practicing scientists that trying to use the scientific literature on its own is not an effective way to develop an accurate picture of the world, especially as one approaches the cutting edge of research. Our experience on the ground working with the literature and with teams trying to replicate and advance it bears this out. In practice, as we say in this post, informal metadata scientists gain from socialization and direct experience is key to using the literature effectively. Taking the literature at face value can, in fact, be harmful. One common manifestation of this is intelligent and well-read amateurs drawing incorrect conclusions because they lack this informal information (e.g. because they don't attend the right conferences).
The bit of data from Longpre et al. that we highlight here is interesting (though, I'd agree that it's not at all conclusive) because it is a hint of evidence that the "well-read amateur" effect may be visible in the context of unsupervised learning in LLMs. My read of figure 5 of that paper (as an amateur in this field-- the irony isn't lost on me) is that dropping some chunks of the scientific literature led to a greater improvement on benchmarks the authors classed as "academic" than dropping anything besides social media, despite the cut literature being 13% of the total training corpus. There are any number of holes one can poke in this, but it's enough to make me want to see a more thorough experiment.
One falsifiable claim we're making here is that as long as the scientific literature is produced within current structures and incentives, AI will need information from outside the literature to make effective use of the literature to predict the world. We don't have insider knowledge of how frontier models are produced, but there seem to be thriving businesses paying human experts decent hourly rates to provide feedback to models. This would be one way of getting that informal information even if the model producers don't realize that's one of the things that's happening when they buy this expert feedback. Alongside this, we'd expect to see continued need for expert human feedback to digest the latest literature as long as valuable parts of the literature come from the current human-driven scientific process. If frontier model training does indeed make extensive use of expert feedback, then the performance of frontier models does not disprove this claim.
If this claim is false, then AI should be able to reliably predict the world from the literature, even at the cutting edge of research. This might show up as accurately estimating the accuracy/replicability of new scientific papers or as inferring the existence of specific failed experiments that were never published based only on the literature. Either of these abilities would be superhuman because humans need to talk to each other often and try things themselves to become skilled at predicting the world from the literature AND the human skill rots if it's not maintained as fields progress. Providing these superhuman abilities for a fee would be an excellent business and so we find the non-existence of this business to be suggestive as well. Even specialist tools like Elicit and Undermind do not claim either of these abilities.
It seems to me that the same flaws in the literature that are blocking LLMs from achieving an effective understanding would also be deleterious to human scientists ability to understand what's real.
> We wonder whether this result holds up in 2026, but we’d bet that it does.
I would be a little more impressed if the result were on something a little larger than a 0.02b-parameter T5 without even any error bars for tiny effects from removing a few % of the total training data. I hesitate to generalize from that to any contemporary LLMs like a Fable at ~10,000b-parameters.
The frontier labs routinely investigate data mixes in extreme detail, and I am unaware of any LLM which doesn't seem to have practically memorized Arxiv, or any reports from, say, the Chinese labs that they have found their own data mix ablations to support the claim that "The Scientific Literature is Poisonous to LLMs".
Dan here, one of the authors of this. Here's an attempt to go deeper. We'd be glad to receive feedback from you or other readers.
As far as we know, it is uncontroversial among practicing scientists that trying to use the scientific literature on its own is not an effective way to develop an accurate picture of the world, especially as one approaches the cutting edge of research. Our experience on the ground working with the literature and with teams trying to replicate and advance it bears this out. In practice, as we say in this post, informal metadata scientists gain from socialization and direct experience is key to using the literature effectively. Taking the literature at face value can, in fact, be harmful. One common manifestation of this is intelligent and well-read amateurs drawing incorrect conclusions because they lack this informal information (e.g. because they don't attend the right conferences).
The bit of data from Longpre et al. that we highlight here is interesting (though, I'd agree that it's not at all conclusive) because it is a hint of evidence that the "well-read amateur" effect may be visible in the context of unsupervised learning in LLMs. My read of figure 5 of that paper (as an amateur in this field-- the irony isn't lost on me) is that dropping some chunks of the scientific literature led to a greater improvement on benchmarks the authors classed as "academic" than dropping anything besides social media, despite the cut literature being 13% of the total training corpus. There are any number of holes one can poke in this, but it's enough to make me want to see a more thorough experiment.
One falsifiable claim we're making here is that as long as the scientific literature is produced within current structures and incentives, AI will need information from outside the literature to make effective use of the literature to predict the world. We don't have insider knowledge of how frontier models are produced, but there seem to be thriving businesses paying human experts decent hourly rates to provide feedback to models. This would be one way of getting that informal information even if the model producers don't realize that's one of the things that's happening when they buy this expert feedback. Alongside this, we'd expect to see continued need for expert human feedback to digest the latest literature as long as valuable parts of the literature come from the current human-driven scientific process. If frontier model training does indeed make extensive use of expert feedback, then the performance of frontier models does not disprove this claim.
If this claim is false, then AI should be able to reliably predict the world from the literature, even at the cutting edge of research. This might show up as accurately estimating the accuracy/replicability of new scientific papers or as inferring the existence of specific failed experiments that were never published based only on the literature. Either of these abilities would be superhuman because humans need to talk to each other often and try things themselves to become skilled at predicting the world from the literature AND the human skill rots if it's not maintained as fields progress. Providing these superhuman abilities for a fee would be an excellent business and so we find the non-existence of this business to be suggestive as well. Even specialist tools like Elicit and Undermind do not claim either of these abilities.
It seems to me that the same flaws in the literature that are blocking LLMs from achieving an effective understanding would also be deleterious to human scientists ability to understand what's real.