Skip to content

Measuring content duplication rate across a portfolio before search does it for you

Content studioUpdated 2026-08-238 min read

In short

Content duplication rate is the share of published pages that sit above a chosen text similarity threshold with at least one other page in the corpus. Compute pairwise resemblance on ten word shingles, strip template boilerplate first, publish the threshold beside the figure, and report it monthly by brand.

A publishing director asks how much of the archive is duplicated. Six brands, about 4,800 live pages, five years of reuse, one editor who has been quietly repointing the same explainer at four audiences since 2023. Nobody in the room can answer, so somebody offers a feeling, and the feeling becomes the number.

Measuring content duplication rate properly takes a weekend of compute and gives you something you can put on the same slide every month. The method has been sitting in the literature since the late nineties and almost nobody in event media runs it.

Why does a pairwise count get expensive so fast?

Duplication is a property of pairs. One page cannot be a duplicate on its own, so the question is always how many other pages it resembles, and that means comparing everything with everything.

For 4,800 pages the number of unordered pairs is 4,800 multiplied by 4,799 and halved, which is 11,517,600 comparisons. That is fine on a laptop if each comparison is cheap. Double the corpus to 9,600 pages and it quadruples to just over 46 million, which is the shape of the problem that made this a research topic in the first place.

The cost only becomes interesting once you decide what a comparison is. Comparing two pages word by word tells you almost nothing, because two pages can share every word in the dictionary and say different things. What you want is shared sequence.

The 1997 method that still holds up

Broder, Glassman, Manasse and Zweig published "Syntactic clustering of the Web" in Computer Networks and ISDN Systems in 1997, and the method they set out is still the one to use.

Their definition is short. "A contiguous subsequence contained in D is called a shingle. Given a document D we define its w-shingling S(D, w) as the set of all unique shingles of size w contained in D." Take the words of a page in order, slide a window of w words along it, and collect every distinct window. A 1,200 word article at w of ten gives 1,191 shingles.

Then compare the sets. Their resemblance measure is the size of the intersection over the size of the union, which is Jaccard similarity computed on shingles instead of on words. Two pages sharing every ten word run score one. Two pages sharing none score zero. A page that has been reworded around a stable skeleton lands in between, and lands there in a way that tracks how much of the original sequence survived.

The scale they worked at is worth knowing, because it settles any argument about whether a portfolio archive is too big for this. They ran it on "a collection of 30,000,000 HTML and text documents from a walk of the web performed by AltaVista in April of 1996", clustered at fifty per cent resemblance, and found 3.6 million clusters covering 12.3 million documents. Their shingle size was ten.

They did not perform 450 trillion comparisons to get there. Each document gets reduced to a fixed size sketch of its shingle set, and resemblance between two documents is estimated from the sketches, which turns an unaffordable exact computation into an affordable approximate one. For a corpus of a few thousand event pages you can skip the approximation entirely and compare the full sets, which is slower per pair and simpler to explain to whoever audits the number.

What do you have to strip out before measuring?

This is the step that decides whether your figure means anything, and it is not in the maths.

Every page on a show site carries navigation, a footer, a sponsor strip, a registration promotion and usually a block of related links. On a typical event CMS that comes to a few hundred words of identical text on every URL in the brand. Two entirely unrelated articles are therefore already duplicates of each other in the boilerplate.

Work it through. Suppose each page renders 1,600 words in total, of which 400 are template. Two unrelated articles share those 400 words and nothing else, so on a naive extraction the intersection is 400 and the union is 1,600 plus 1,600 minus 400, which is 2,800. That is a similarity of 400 over 2,800, about 14 per cent, before either article has said a thing. Every pair in the corpus gets that floor, the range compresses into the top of the scale, and a threshold you chose on intuition now sits in the wrong place.

Extract the article body only. Most CMS platforms will give you the field directly through an API, which is better than parsing rendered HTML. Where you have to parse, strip by selector and then check ten pages by hand against the extracted text, because a silent extraction bug produces a duplication rate that is confidently wrong and looks plausible.

Two more decisions belong here. Lowercase everything and collapse whitespace, or Kathryn on one page and KATHRYN on another read as different shingles. And decide what to do with quoted material, because a panel write up that quotes 200 words of a speaker will legitimately match every other piece quoting the same speaker. Excluding blockquotes is defensible. Including them is defensible. Doing it differently in March and April is not.

Working the rate on a corpus of 4,800 pages

Run it once across the whole archive to size the debt, then decide what to publish monthly.

Take the 4,800 pages, strip to body text, shingle at ten, compute resemblance for all 11,517,600 pairs. Now count pages, and count them once each: a page belongs in the numerator if it exceeds the threshold against at least one other page, however many partners it has.

Suppose at a threshold of 0.7 that comes to 612 pages. The rate is 612 divided by 4,800, or 12.8 per cent. At 0.5, the threshold Broder and colleagues used, 1,344 pages qualify, which is 28 per cent. At 0.9, where the pages are close to identical, 118 qualify, which is 2.5 per cent.

Three numbers from one run, and the useful move is to report all three rather than argue about which is correct. The 2.5 per cent is your straightforward duplicate problem and it is small. The 28 per cent is the size of the reuse habit. The gap between them is where the editorial judgement lives.

Then split by brand, because the portfolio average hides the thing you can act on. If brand D contributes 214 of the 612 pages above 0.7 while publishing only 640 pages of the 4,800, its own rate is 214 over 640, or 33 per cent, against a portfolio figure of 12.8 per cent. One desk is generating most of the duplication, and now you know which one and can go and ask why.

What Google actually penalises, and what it does not

The search argument for this measurement gets overstated in both directions.

Google's spam policies, current in 2026, define scaled content abuse as "when many pages are generated for the primary purpose of manipulating search rankings and not helping users", and the scraping section gives as an example taking content and modifying it "only slightly (for example, by substituting synonyms or using automated techniques)" before republishing. A portfolio with a 12.8 per cent duplication rate produced by human editors reusing their own work is a long way from that description.

What a high rate does reliably predict is that your pages compete with each other, which is a different failure with different symptoms and its own diagnosis. A duplication rate tells you where to look. It does not tell you that anything is losing traffic, and reporting it as if it did will get the whole measurement dismissed the first time somebody checks a high scoring pair and finds both pages ranking fine.

Keep the two separate in the reporting. This number sizes the corpus. Whether a given pair is costing you clicks is a question for query data.

Where this stops

Text similarity is blind to meaning. Two pages can score 0.15 against each other and say exactly the same thing in different words, which is the normal output when two editors cover the same announcement independently, and no shingle based method will ever catch it. If that case matters to you, embeddings are the tool, and they bring their own calibration problem because the similarity scores are not interpretable in the way a Jaccard figure is.

The measure is also unstable on short pages. Broder and colleagues note that shingling works less well on very short documents, and the reason is mechanical: a 200 word page at w of ten has 191 shingles, so a handful of shared runs moves the score a long way. Set a minimum length, report it, and put everything below it in a separate bucket rather than pretending the score is comparable.

Finally, this is a corpus measurement and it says nothing about records. A duplicate rate over a registration file counts rows that describe the same person and uses a completely different denominator, which is a data quality metric with its own definitional traps and no relation to this one beyond the word. Two numbers, two owners, and putting them on the same dashboard confuses everybody.

Pick one brand, export the body text of everything it has published in the last two years, and run the pairwise comparison at 0.5, 0.7 and 0.9. It is an afternoon of work and it gives you a curve for one desk. That curve is enough to decide whether you need a gate before publication or a routing decision on the pairs you already have, and whoever runs the content operation can read it without a briefing.

Questions people ask about measuring content duplication rate

How do you measure duplicate content across a whole website?
Break each page into overlapping runs of consecutive words, usually ten, and compare the resulting sets between pages. The overlap between two sets divided by their union gives a similarity score between zero and one. Count the pages that exceed a chosen score against any other page, and divide by the corpus size.
What similarity threshold counts as duplicate content?
No published threshold applies to every corpus, so the honest approach is to report the curve. Broder and colleagues clustered the web at fifty per cent resemblance in 1997. Event portfolios usually want something stricter for the headline figure, around seventy per cent, with the looser bands shown underneath it.
Why does template boilerplate distort a duplication score?
Navigation, footers, sponsor blocks and event promotion repeat on every page, so two entirely unrelated articles already share hundreds of identical words before either says anything. That shared text raises every pair's score by roughly the same amount, which compresses the range and makes the threshold meaningless.

Related reading

All content and media articles