4 Comments

User's avatar
gwern's avatar

> We wonder whether this result holds up in 2026, but we’d bet that it does.

I would be a little more impressed if the result were on something a little larger than a 0.02b-parameter T5 without even any error bars for tiny effects from removing a few % of the total training data. I hesitate to generalize from that to any contemporary LLMs like a Fable at ~10,000b-parameters.

The frontier labs routinely investigate data mixes in extreme detail, and I am unaware of any LLM which doesn't seem to have practically memorized Arxiv, or any reports from, say, the Chinese labs that they have found their own data mix ablations to support the claim that "The Scientific Literature is Poisonous to LLMs".

Indrajeet Yadav's avatar

Because LLMs are probabilistic machines that predict the next set of tokens (words) based on patterns they’ve learned to recognize from their training data, their output depends on data quality.

The irony, as Bob Garvin (Donald Sutherland) notes in the 1994 film Disclosure, “. . . that’s the legacy of the modern age. We have information but no truth . . .”

Scientific literature is no exception: “Once authorship and citations started being used as a metric, all the signal was gamed out of them [scientific metadata].”

Goodhart’s Law reminds us that when a measure becomes a target, it ceases to be a good measure. Part of the problem is blind faith in metrics. Another is intangibility. A single metric can hardly capture a broad, complex phenomenon.

Speaking of intangibility, journalists rely on the old adage that the [whole] truth lies somewhere in between [competing claims]. Finding it requires nuance and emotional intelligence, qualities that are uniquely human. At least for now.

But then, even scientists with good emotional intelligence struggle to tread through the “minefield of half-truths, convenient omissions, and outright lies that can’t be distinguished from the innocent facts and findings around them” without the benefit of socialization and scientific metadata.

Why then, expect an LLM, with little or no nuance, to reliably separate music from noise? When the noise is its training data?

2 more comments...

No posts

Ready for more?