How Do You Stop an AI Content Pipeline Repeating Itself?
How do you stop an AI content pipeline repeating itself?
Compare every new topic against everything you have already published, before you write it. The check that works is numeric rather than editorial. You turn each title and outline into an embedding, measure similarity against your existing library, and refuse anything above a threshold you set deliberately.
We think about this constantly, because this blog publishes on a daily cadence. At that pace, memory is not a reliable defence. Nobody can hold hundreds of titles in their head and reliably notice that today's idea is last month's idea wearing a different heading.
The failure is quiet, too. You do not get an error. You get two articles that both half answer the same question, and neither one ranks.
Why do AI pipelines produce near duplicates?
Because a model generating topics reaches for the same strong ideas every time. Ask for ten article topics in a niche and you get the ten most obvious ones. Ask again next week and you get eight of the same, phrased differently. The model has no memory of what you already shipped.
Topic space also narrows as you publish. The first hundred articles in a category cover the broad questions. By the third hundred, every genuinely new topic sits closer to something existing, so the gap between fresh and repetitive gets thinner.
Human editors drift the same way, just slower. The pipeline only makes the problem visible sooner, which is a favour disguised as a defect.
What counts as too similar?
Two articles are too similar when a reader who found one would gain nothing from the other. That is the standard worth holding, and it is stricter than avoiding copied sentences. Different words answering the same question for the same reader is duplication, whatever a plagiarism checker says.
The useful distinction is angle. An article about choosing a tool and an article about migrating off that tool serve different moments, even though they share most of their vocabulary. Two articles about choosing the tool, one framed as a question and one as a comparison, do not.
So the test is not surface overlap. It is whether the second piece has something the first did not: a different decision, a different reader, or genuinely new information.
How do you actually measure similarity?
With embeddings and cosine similarity. You convert each piece of text into a vector of numbers, then measure the angle between vectors. Text about the same thing lands close together, regardless of wording, which is exactly the property you need for catching rephrased repeats.
OpenAI's embeddings guide names cosine similarity as its recommended distance function and notes that "OpenAI embeddings are normalized to length 1." That normalisation matters in practice, because it means you can compare scores directly without extra maths.
The guide also lists the jobs embeddings are built for, and two are exactly this task. It describes clustering as where "text strings are grouped by similarity" and diversity measurement as where "similarity distributions are analyzed." You are doing both: grouping near neighbours and watching whether your library is getting narrower over time.
Which embedding model should you use for this?
The cheap one, almost certainly. OpenAI's guide lists text-embedding-3-small at 1536 dimensions and text-embedding-3-large at 3072 dimensions, with both accepting a max input of 8192 tokens. For duplicate detection, the smaller model is generally enough, because you are looking for obvious neighbours rather than fine distinctions.
Cost makes the case plainly. The guide gives text-embedding-3-small at 62,500 pages per dollar against 9,615 pages per dollar for text-embedding-3-large. That is close to a sixfold difference for a job where precision at the margins does not change your decision.
Embedding your whole archive is a one time cost, and then each new candidate is a single call. At those rates the running cost of this safeguard rounds to nothing, which removes the usual excuse for skipping it. We compared these models for a different purpose in our piece on choosing an embedding model for site search.
Where in the pipeline should the check run?
At topic selection, before a single word is written. This is the whole point. Catching a duplicate after drafting means you have paid for the article and now have to throw it away, which creates pressure to publish it anyway.
We run it against a stored list of everything already published, including the title and a short description of the angle. The candidate topic gets embedded the same way and compared against every entry. Anything too close is rejected and replaced from a reserve list before writing begins.
A second checkpoint after drafting is worth having, but it should almost never fire. If it fires often, your topic selection is the problem and you are papering over it at the wrong end.
Keep the stored list somewhere boring and durable. A table in your database, a JSON file in the repository, an Airtable base. What matters is that it is written to the moment something publishes, not reconstructed later from the live site.
What threshold should you set?
Lower than feels comfortable, then tune it with real rejections. Start strict, look at what gets blocked, and loosen only in the cases where you genuinely disagree with the machine. Starting loose and tightening later is the common mistake, because by the time you notice, the duplicates are already published and indexed.
Titles alone are a weak signal, because two very different articles can share a phrasing pattern. Comparing the title plus a sentence of angle works far better, since the angle is where genuine difference lives.
Expect to review borderline cases by hand at first. After a few dozen decisions you will find the score where your judgement and the number agree, and that is your threshold. There is no universal figure, because it depends on how narrow your niche is and how much you have already published.
What do you do when a topic fails the check?
Either sharpen it into something genuinely new or drop it and take the next one. The lazy move is to reframe the heading and proceed, which defeats the check entirely. If the embedding says it is the same article, changing the title does not make it a different article.
Sharpening means finding the narrower question inside the broad one. If a general guide already exists, the new piece can take one decision from it and go deeper than the guide had room for. That is a real article with a real reason to exist.
Sometimes the right answer is to improve the existing piece instead. A near duplicate is often a signal that the original was incomplete, and updating it beats splitting the topic across two weaker pages. Our notes on fixing keyword cannibalisation cover what happens when you do not.
Does Google punish near duplicates?
Not as a penalty, but near duplicates compete with each other, and at volume they raise a real risk. Google treats duplication as something to consolidate rather than punish. Its documentation says that without a canonical, "Google will identify which version of the URL is objectively the best version to show to users in Search."
The sharper risk sits in the spam policies. Google defines scaled content abuse as when "many pages are generated for the primary purpose of manipulating search rankings and not helping users," and names using generative AI "to generate numerous pages without user value" as an example. A pipeline producing near duplicates at volume is walking toward that description.
So the similarity check is not only an editorial nicety. It is the mechanism that keeps a high volume operation on the right side of a published policy, and it is the thing you would want to be able to point at if anyone asked.
What does this look like running well?
Boring. A topic gets proposed, scored, and either accepted or swapped for the next one, and nobody has to remember anything. The interesting output is the rejection log, because it tells you where your coverage has become saturated and which category needs genuinely new thinking.
The deeper benefit is that it forces the question every publishing operation should ask anyway. Does this piece have something the last one did not. A machine asking that on every article is more consistent than a person asking it sometimes. We wrote about the surrounding system in our piece on automating blog publishing.
If you are building a content engine and want a second opinion on where the guardrails belong, we are happy to walk through ours with you. You can find us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.