I measured 7,872 articles and got a clean answer: long headlines get read less. Then I split the same numbers by outlet and the answer reversed completely. This is statistics' most famous trap, and I walked straight into it.
I asked a simple question: does headline length affect how much an article gets read? Our archive holds 7,872 articles from the last 60 days, each with its headline and its read count. That is an ideal dataset for answering it.
The answer came back clean. Headlines between 50 and 89 characters averaged 20.6–21.0 reads. Anything over 90 characters averaged just 17.0–17.6. Long headlines do roughly 20% worse. Editorial rule written: keep headlines short.
One thing bothered me. The number was too tidy. So I asked the same question inside a single outlet — because the BBC and Yahoo Finance do not write headlines the same way.
Inside the BBC the picture inverted: headlines under 50 characters averaged 18.8, those of 70–89 averaged 23.7, and those over 90 averaged 29.0. The Guardian pointed the same way: 17.0 for short, 24.2 for long. In both newsrooms, longer headlines did better.
How can long headlines be the worst in the combined data and the best inside every outlet? The answer is that I was not measuring length. I was measuring newsrooms.
Sources like Yahoo Finance file large volumes of short, near-identical market headlines, and their average read count is low. The BBC and the Guardian write longer, more explanatory headlines and are read more in general. In the pooled table, 'long headline' really meant 'BBC and Guardian', and 'short headline' meant 'market wire'.
Statistics has a name for this: Simpson's paradox. A trend that appears within groups can reverse when the groups are combined. You read about it in a textbook; seeing it in your own data is a different feeling entirely.
There are two takeaways. The narrow one: 'shorten the headline' would have been the wrong rule for us. The broader one: holding 30 outlets in one table is both the strength of this project and its trap. Averages are easy to take, and an average across mixed sources often measures the source mix rather than the question.
I am writing this up precisely because the result is not flattering. Had I published the first number, nobody would have checked it — it looked convincing and the chart was handsome. The error only surfaced because I asked myself one more question: what if the outlets differ?
Next time you see a site claim that 'research shows shorter headlines perform better', ask one thing: does that hold inside individual publications, or only after everything was poured into one bucket?