How Do You Stop an Automation Creating Duplicates?
Why does your automation keep creating duplicate records?
Because almost every system that calls your automation is willing to call it twice. Stripe's webhook documentation states it plainly: webhook endpoints might occasionally receive the same event more than once. Your automation did not misfire. It was asked to run twice and had no way to notice.
This is the single most common failure we see in automation work, and it rarely looks like a failure. Nothing errors. No alert fires. There are just two of something, and someone in sales notices a week later.
The fix is an old idea from payments engineering called idempotency, and it transfers neatly to any automation that writes data. Here is how it works and how to retrofit it when the tools you are using do not offer it.
What is an idempotency key?
A unique string the caller attaches to a request so the server can recognise a retry. Stripe describes it as a key that a client generates, which the server uses to recognise subsequent retries of the same request. Send the same key twice and the second call returns the first call's result instead of doing the work again.
The important part is that the caller generates the key, not the server. That seems backwards until you think about the failure case. If your connection drops before the response arrives, you do not know whether the write happened. Only you can name the attempt in a way that survives your own uncertainty.
So an idempotency key is really a name for an intention. It says: this is the one time I meant to create this invoice, and any repeat of this message is the same intention, not a new one.
How does Stripe's idempotency actually behave?
It caches the outcome, including failures. Stripe's documentation says its idempotency works by saving the resulting status code and body of the first request made for any given key, regardless of whether it succeeds or fails, and that subsequent requests with the same key return the same result, including 500 errors.
That last detail surprises people. A cached 500 is not a bug. It means the server is refusing to guess whether your original attempt partly succeeded. If you want a genuinely fresh attempt, you need a fresh key, and you need to have decided that a retry is safe.
Stripe also guards against careless reuse. Its documentation says the idempotency layer compares incoming parameters to those of the original request and errors if they are not the same, to prevent accidental misuse. And keys are not kept forever: Stripe says keys can be removed from the system automatically after they are at least 24 hours old, and a reused key after pruning generates a new request.
Is there a standard for the Idempotency-Key header?
There is a draft, and it has lapsed. The IETF HTTPAPI working group document draft-ietf-httpapi-idempotency-key-header-07 describes a header that can be used to make non-idempotent HTTP methods such as POST or PATCH fault-tolerant. It was published on 15 October 2025 and expired on 18 April 2026 without becoming an RFC.
The draft is still useful as a description of the shared convention. It says the idempotency key must be unique and must not be reused with another request with a different payload, and that if someone does reuse a key with a different payload, the resource should reply with HTTP 422.
What that expiry means in practice is that you cannot assume any given API implements the header, or implements it the same way. Read the vendor's own documentation each time. The pattern is widely shared, the specification is not binding, and the details differ.
What do you do when the API you call has no idempotency support?
You build the check on your side. Keep a small table of operations you have already performed, keyed by something stable from the incoming event, and consult it before you write. This is the same mechanism, moved one step closer to you.
Stripe recommends exactly this shape for its own webhooks. Its best practices say you can guard against duplicated event receipts by logging the event IDs you have processed and then not processing already-logged events. The event ID is the key, and your own log is the cache.
The ordering detail matters too. Stripe says it does not guarantee delivery of events in the order they were generated, and warns against using the created timestamp to determine event order or whether you have already processed an event, advising you to track event IDs instead. An automation that assumes order will eventually process a cancellation before the thing it cancels. Our notes on webhooks versus polling go further into that trade-off.
Why can you not just check whether the record already exists?
Because two copies of your automation can check at the same moment. Both read, both see nothing, both write. The check and the write have to be one operation, not two, or the race is still there and just narrower.
In a database this means a unique constraint on the key, and letting the second insert fail harmlessly rather than asking politely first. In an automation platform it usually means a single upsert step rather than a search step followed by a create step. If your tool only offers search then create, you have a race you cannot close inside that tool.
Stripe hints at the same subtlety from the server side. Its documentation notes that results are only saved after execution of an endpoint begins, and that a request conflicting with another request executing concurrently does not get a saved idempotent result. Concurrency is a separate problem from retries, and both need an answer.
How do automation platforms handle duplicates for you?
Partly, and with assumptions you should know about. Zapier's platform documentation says it makes an initial call to your API, then caches and stores each id field in its database, comparing incoming IDs against the cached ones to spot new items. That is deduplication, done for you, for polling triggers.
The assumptions are where it gets interesting. Zapier's documentation notes that when a Zap is turned off, that list is cleared. It also addresses updated items by recommending you set the id to a new combined value that is unique for every update, and it expects reverse-chronological sorting so new items appear first.
So the platform protects you from reprocessing the same unchanged item, which is genuinely useful. It does not protect you from the automation running twice because someone duplicated the workflow, or from a downstream step failing after an earlier step already wrote. The guarantee covers the trigger, not the whole chain.
What makes a good idempotency key?
Something stable, unique, and meaningless. Stripe suggests V4 UUIDs or another random string with enough entropy to avoid collisions, says keys can be up to 255 characters long, and advises against using sensitive data such as email addresses or personal identifiers as keys.
That last warning is worth repeating because it is tempting to break. An email address as a key feels convenient and stable. It also puts personal data into log lines, cache entries, and error messages in systems that were never designed to hold it. Use an opaque value and keep the mapping somewhere appropriate.
When the event comes from someone else's system, the best key is usually the one they already assigned: the event ID, the message ID, the row ID plus a version. When you are the originator, generate a UUID at the moment the intention is formed, and carry it through every retry of that intention rather than minting a new one per attempt.
How do you test for duplicates before they reach production?
Replay the same input twice and check the result. That is the whole test, and almost nobody runs it. Take a real event payload, send it to the automation, then send the identical payload again and count the rows. If you get two, you have found the bug on a Tuesday instead of at quarter end.
Then test the harder case: send two copies at the same moment rather than in sequence. This is where search-then-create setups fall over, and it is the case that production will eventually produce when a provider retries aggressively after a timeout.
Build the retry window into the test too. Stripe says it attempts delivery for up to three days with exponential back off in live mode, and that events can be manually resent for up to 15 days from the Dashboard or 30 days with its CLI. An automation that only dedupes against the last hour of history is not deduping against reality. Our piece on testing automations without real data covers how to do this safely.
How do we handle this in our own automation work?
We treat every write as something that might arrive twice, because it will. The rule we apply is that any step which creates or changes a record needs a key and a constraint, and if the target system cannot enforce uniqueness, we keep our own ledger of what we have already done.
We also prefer loud duplicates over silent ones. A unique constraint that rejects the second write leaves a clean error in a log. A workflow that quietly creates two records leaves nothing at all, which is why duplicate bugs tend to be discovered by a human rather than by monitoring. Our notes on silent automation failures come from the same instinct.
Where we are honest about the cost: this adds a table and a little bookkeeping to every automation that writes data. It is not free, and on a workflow that posts a message to a channel nobody reads, it is probably not worth it. On anything touching a CRM, an invoice, or a customer-facing record, it is the cheapest insurance available.
Where does idempotency go as automations get more agentic?
It gets more important and harder to retrofit. A scripted workflow repeats the same call. An agent deciding its own steps may attempt an equivalent action by a slightly different route, which an exact-payload comparison will not catch. The key then has to name the intention at a higher level than the request.
Our working view is that the teams who already put keys and constraints on their writes will adapt to agent-driven automation with far less pain than the teams relying on step order and good luck. The discipline is the same, the stakes are just higher when the caller is improvising. Our notes on handling automation failure cover the surrounding habits.
If you have an automation quietly producing duplicates and you want a second opinion on where to put the key, we are happy to look at it with you. Find us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.