Articles · GDPR

2026-12-22 12 min EN / FR

GDPR and test environments: stop copying production data

The most common GDPR problem in small companies is not in the production system, it is in the copies. A developer restores last night's backup to a staging server "just to reproduce a bug", and now real customer data sits on a machine with weaker passwords, no monitoring, and a dozen people who can log in. Nobody decided this. It just happened.

This article is practical guidance, not legal advice. For your specific situation, check with your data protection officer, a lawyer, or the CNIL's published guidance.

The core problem

Under the GDPR, personal data does not stop being personal data because it is in a test environment. The same principles apply:

  • Purpose limitation: data collected to fulfil orders was not collected so developers could test features.
  • Data minimisation: you should use only the data you need. Testing a checkout flow rarely needs 40,000 real customers.
  • Security: you must protect it appropriately, wherever it sits.
  • Storage limitation: copies should not live forever.
  • Accountability: you must be able to show how you handle it.

A staging server with a full production copy usually fails on most of these at once.

Step 1: Find where production data already leaks

Before fixing anything, look. Common places real data hides:

  • Staging and test databases restored from backups
  • Developer laptops, with local database dumps in download folders
  • Shared drives and chat messages where someone shared an export or a screenshot
  • Logs, which often contain emails, names, IP addresses and request bodies
  • Error tracking and monitoring tools, which capture the data attached to each error
  • Old backups and snapshots of test servers
  • Analytics and support tools receiving events from a staging site that is wired to the real account
  • Spreadsheets created for "a quick analysis"
  • Third-party contractors' machines and accounts

List each place, who has access, and since when. This inventory is also what you need for your processing register. Pair it with a quick account ownership audit so you know which logins and tools actually reach each environment.

Step 2: Choose the right type of test data

You have four options, from safest to riskiest.

  1. Synthetic data (generated from scratch). Fake customers, orders and products generated by a script or library. Nothing links to a real person. This is the best default for development and automated tests.
  2. Anonymised data. Real data transformed so that people can no longer be identified, by anyone, by any reasonably available means. If it is truly anonymous, it falls outside the GDPR. This is a high bar and harder to reach than most people think (see below).
  3. Pseudonymised data. Identifiers replaced by codes or fake values, but re-identification is still possible, for example because the original is kept or because fields combine to identify someone. This is still personal data under the GDPR. It reduces risk but does not remove your obligations.
  4. Real production data. Only when strictly necessary, for a limited time, in a controlled environment with production-level security, and with a documented justification.

Step 3: Why "anonymised" is harder than it sounds

Replacing names with random ones is not enough. People can often be identified by combining remaining fields:

  • A rare job title + a small town + a birth date
  • A delivery address, even without a name
  • Free-text fields: support messages, notes and comments contain names, phone numbers and health details
  • Attachments: PDFs, invoices and images
  • Unique behaviours: an order history that only one person has
  • Timestamps and IDs that can be matched to other sources
  • Rare values: the only customer in a country, or the only account over a certain size

A practical test: could someone with access to this data, plus other information they could plausibly obtain, single out an individual? If yes, treat it as personal data.

Do not call data "anonymised" in your documentation unless you have actually assessed this. "Pseudonymised" is often the honest label.

Step 4: Build a safe copy process, if you need real structure

Sometimes you need data shaped like production: volume, odd cases, edge cases. A safe pipeline:

  1. Restore a production backup into a locked-down, temporary environment (never straight to staging).
  2. Run a masking script that transforms the data before anyone else can access it.
  3. Verify the result with automated checks (see below).
  4. Publish the masked database to the test environments.
  5. Destroy the temporary copy.

What the masking script should do:

  • Replace direct identifiers: names, emails, phone numbers, addresses, national IDs, bank details
  • Use realistic fakes with a proper format, so validation and forms still work
  • Keep referential integrity: the same customer must get the same fake identity everywhere. Deterministic generation (for example, deriving the fake value from a secret-keyed hash of the original ID) does this, but keep the secret protected, since weak hashing of small value sets can be reversed
  • Scramble or clear free text and drop attachments
  • Remove secrets: API keys, tokens, webhook URLs and password hashes
  • Generalise sensitive fields such as birth dates (keep the year, or shift by a random offset) and exact locations
  • Drop what tests do not need: old records, whole tables of sensitive data, archived history

Make the script a maintained part of your codebase. When someone adds a column with personal data, a rule in the checklist asks whether the masking script needs updating.

Review your test environments

We map where production data leaks, set up masking or synthetic data, and lock down staging without blocking your team.

Step 5: Neutralise the side effects

Test environments can act on real data. A copied database may contain real email addresses, phone numbers and live integration settings. Prevent:

  • Real emails being sent. Route all outgoing mail to a catch-all tool (like MailHog or Mailpit) or an allowlist, and double-check that the setting cannot be accidentally turned off.
  • Real SMS and notifications.
  • Real payments. Use the provider's test mode and test keys only.
  • Webhooks calling production systems, or marketing tools receiving fake signups.
  • Scheduled jobs (reminders, invoices, newsletters) running against copied data.
  • Analytics and tracking polluting your real dashboards.

Replace every production key and URL in the test configuration. A common rule: the test environment should be unable to talk to production credentials.

Step 6: Secure what remains

If, after all this, a test environment still contains real personal data, then it must be protected like production:

  • Access limited to named people who need it, with strong authentication
  • Not publicly reachable: no open staging URLs indexed by search engines
  • Encryption at rest and in transit
  • Logs and monitoring
  • Documented retention: a deletion date
  • Backups handled with the same care
  • Included in your processing register and security review

"It is only staging" is not a defence. Several well-known breaches started with an exposed test server.

Step 7: Do not forget your suppliers and developers

  • External developers and agencies who access personal data on your behalf are usually your processors, and you need a written agreement setting out what they can do (a data processing agreement under Article 28). Document expectations in your exit plan as well.
  • Tools that receive your data (error tracking, logging platforms, support tools, AI services) are processors too. Check where the data goes, including transfers outside the EU, and configure scrubbing of personal data where possible.
  • Contractors' own laptops: define what may be stored on them, and require deletion at the end of the engagement.
  • Offboarding: remove access to test databases and dumps when people leave.

Step 8: Prepare for the day it goes wrong

If personal data in any environment is exposed, you may have to notify the supervisory authority (the CNIL in France) within 72 hours of becoming aware, when the breach is likely to create a risk to individuals, and possibly inform the people concerned. A test environment breach counts if real data was in it. Know in advance:

  • Who decides whether it is a reportable breach
  • Who contacts the authority
  • Where your inventory and logs are, so you can say what was exposed

Have this written down before you need it.

What to put in your documentation

  • Your processing register should mention test and development uses if real data is involved
  • A short policy: "Production data is not copied to non-production environments except under procedure X"
  • The masking procedure, with a date and an owner
  • The list of environments, who has access, and what data each contains
  • Retention periods for copies
  • For new projects, a data protection impact assessment where the risk level requires it

Common mistakes

  • Restoring a production backup to staging for convenience
  • Calling pseudonymised data "anonymous"
  • Leaving free-text fields and attachments untouched
  • Forgetting logs, error trackers and analytics
  • Test servers publicly accessible and indexed
  • Real SMTP settings left in the test configuration
  • Developers keeping local dumps indefinitely
  • No agreement with the contractors who handle the data
  • No owner for the masking script, so it quietly goes out of date

Quick audit

  • I know every environment and what data it contains
  • No test environment holds real personal data without a written justification
  • A masking or synthetic-data process exists and is maintained
  • Test environments cannot send real emails, SMS or payments
  • Free text, attachments and secrets are handled in the masking
  • Logs and error tracking do not capture personal data unnecessarily
  • Contractors and tools handling data are covered by agreements
  • Access to test data is limited, authenticated and reviewed
  • Copies have a retention period and get deleted
  • We know who to call and what to do within 72 hours of a breach

Related

From accidental copies to a clear test-data policy

If staging already holds production data and you want a safe pipeline your team can keep using, write us.