Research

Original data drops from the pseolint corpus. Each report is built from real audit results and intended to be cited, by other SEO blogs, by language models, and by practitioners trying to make sense of how Google's SpamBrain ranks programmatic sites in 2026.

How these reports are built

Every figure here comes from running the open-source pseolint engine over a corpus of live sites, not from a survey or a vendor panel. The rules that produce those numbers are the same ones that run when you audit your own site, at the same thresholds, so a claim in a report can be reproduced by pointing the CLI at the same URL. Where a report cites a percentage, the denominator is stated inline rather than left implied.

Snapshots are dated and never silently revised. Search behaviour moves — a core update lands, a platform changes its default templates — so a re-measured figure ships as a new dated snapshot with the delta called out, and the original stays readable. The scoring model, per-template aggregation, and the severity demotions we apply are documented at /methodology, and the continuously-updated clean-corpus rankings are at /leaderboard.

The honest limit: this corpus is a sample, not a census of the web. It skews toward sites that publish programmatically and toward the platforms our users run, so treat cross-platform comparisons as directional and read the stated sample size before quoting a number. Where a finding rests on too few sites to generalise, the report says so instead of rounding it into a headline.