How Do You Test an Automation Without Real Customer Data?
How Do You Test an Automation Without Real Customer Data?
With a small set of handmade records that cover your awkward cases, held in a sandbox the automation cannot escape from. Copying production data into a test environment feels faster on Monday and becomes a problem you cannot undo on Friday.
Almost every automation we are asked to look at was tested against live data at some point. Usually once, usually by the person who built it, usually with the reasoning that test data would not be realistic enough.
That reasoning is half right. Realism matters. The answer is to build realistic fake data rather than to borrow real data, and it is far less work than it sounds.
What Is Actually Wrong With Using a Copy of Production?
You have taken data people gave you for one purpose and used it for another, in an environment with weaker controls and more people looking at it. Under the GDPR that is a problem in two separate places.
Article 5(1)(b) requires that personal data be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes." Your customer did not hand over their address so it could seed a test run.
Article 5(1)(c) requires data to be "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed." A full copy of your customer table is not the minimum necessary to check that a workflow fires correctly.
Article 5(1)(e) adds a third angle, requiring data be "kept in a form which permits identification of data subjects for no longer than is necessary." Test databases are famously immortal. Nobody deletes the one from the migration two years ago.
Does the Regulation Say What to Do Instead?
It points at the technique directly. Article 25(1) asks controllers to implement measures "such as pseudonymisation" in order to give effect to principles including "data minimisation." Article 25(2) requires that "by default, only personal data which are necessary for each specific purpose of the processing are processed," across the amount collected, the extent of processing, the storage period and who can reach it.
That is a fair description of a good test environment. Only what the test needs. Only for as long as the test needs it. Only reachable by the people running it.
You do not need a legal opinion to act on this. You need a decision that production data does not leave production, and then a bit of engineering to make that decision easy to keep.
What Should Your Test Records Actually Contain?
The cases that break things, not the cases that work. A test set of twenty carefully chosen records is worth more than twenty thousand borrowed ones, because the borrowed ones are mostly identical and none were chosen on purpose.
Build for the edges. A name with an apostrophe in it. A name with no surname. An address in a country with no postcode. A company field left empty. A very long email address. A record with two phone numbers in one field, because someone typed it that way. A duplicate of another record with different capitalisation.
Then add the shape problems: a record missing the field your automation keys on, a date in the wrong format, a currency amount with a comma in it, and a record that is legitimately identical to one already processed.
Every one of those is a real failure mode we have hit. None of them requires a real customer.
Where Do You Get Realistic Fake Data?
Generate it. The libraries for this are mature and free, and they produce names, addresses, phone numbers and company names that look plausible in the locale you ask for. Faker for Python and PHP, and the Faker family of ports for JavaScript and Ruby, all do this well.
For volume testing, generate more of the same rather than borrowing. If you need fifty thousand rows to see whether your rate limiting holds up, fifty thousand generated rows test that just as well as real ones, and cost you nothing if the test database leaks.
The one thing to avoid is a half measure. Taking production data and replacing the names, while keeping the addresses, order histories and support tickets, does not give you anonymous data. It gives you a puzzle with most of the pieces still in the box.
What Does a Proper Sandbox Look Like?
An isolated environment with its own credentials, where the actions the automation takes cannot reach a real person. The good platforms provide this and document it clearly.
Stripe's documentation is a useful model. It describes a sandbox as "an isolated test environment" you can use "to test Stripe functionality in your account, and experiment with new features without affecting your live integration," and notes that "when testing in a sandbox, the payments you create aren't processed by card networks or payment providers."
It also makes a recommendation worth copying wherever you can: "Use separate sandboxes for local development and continuous integration (CI) so automated tests don't affect your settings or data." One shared test environment for everything is how a colleague's experiment ends up overwriting the state your test depended on.
Stripe is honest about the gaps too, listing limitations such as not being able to create connections between a Connect platform's sandbox and connected account sandboxes. Every sandbox has edges. Knowing them beats discovering them.
What About Tools With No Sandbox at All?
Many of the tools in a small team's stack have none, and that is where most accidents happen. A spreadsheet, an email tool, a scheduling app: no test mode, one environment, live consequences.
Two things make this survivable. First, create a parallel container inside the live tool: a separate base, a separate workspace, a separate list, named unmistakably as test. Second, point every outbound action at addresses you own. A test email that reaches a real customer is the failure this whole exercise exists to prevent.
We also add a safety net at the boundary. Before any step that sends something to a person, check the recipient against an allowlist when the automation is running in test mode, and refuse everything else. It takes ten minutes and it has saved us more than once.
How Do You Stop Test Runs From Reaching Real People?
Make the mode explicit, and make the unsafe mode the one you have to ask for. An automation should know whether it is running in test or live, and that flag should come from the environment rather than from a variable someone might forget to flip back.
Then make the difference visible. Test runs write to a test destination, use a test sender address, and tag every record they touch. When someone asks whether Tuesday's run was real, you should be able to answer from the data rather than from memory.
The related discipline is knowing when the automation has gone wrong at all, which is a separate problem we cover in stopping an automation from failing silently. Test mode protects your customers. Monitoring protects your data.
Should Every Automation Get This Treatment?
No, and pretending otherwise is why teams skip it entirely. Match the effort to what the automation can destroy.
An automation that reads data and writes a summary into a document needs very little. An automation that sends email to customers, writes to your CRM, changes a subscription, or moves money needs all of it. The question to ask is what a bad run would cost to undo, and whether it could be undone at all.
That framing also decides where your test data effort goes. Spend it on the handful of automations that touch people and money, and keep the rest simple. Our notes on customer data in AI automations make a similar argument about what deserves real controls.
What Is the Smallest Version of This You Could Do Today?
Three things, in an afternoon. Create one test container in each tool your automation touches and name it clearly. Hand write fifteen records that cover the awkward cases listed above and commit them to your repository so they are version controlled and reviewable. Add an environment flag with an allowlist on anything that sends.
That is not a full test strategy. It is enough to stop the two failures that actually happen: a test run reaching a customer, and a production export sitting forgotten in a staging database. The rest can grow as the automation does. If your automations are further along than your testing is, and you would rather not find out the hard way, we are glad to talk it through with you at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.