Articles · Automation

2027-01-20 12 min EN / FR

Syncing two systems: the five ways it breaks

"Just sync the shop with the ERP" sounds like a small task. It is one of the most common sources of long-running bugs in small companies. Stock that is wrong by three units, an invoice created twice, a customer updated in one place and overwritten from the other. The reasons are surprisingly consistent.

The five failure modes below come up again and again. Design against each of them.

First, decide who is the boss of each piece of data

Before writing a line of code, answer this for every data type (customers, products, prices, stock, orders, invoices):

Which system is the source of truth?

If both systems can edit the same field, you will have conflicts. The cleanest design gives each field one owner:

  • Products and prices: owned by the ERP (see also Odoo or custom software? How to decide)
  • Orders: created in the shop, then owned by the ERP once imported
  • Customer addresses: owned by whichever system the customer edits, with a clear rule

A two-way sync is possible, but it is much harder. Prefer one-way flows wherever you can.

Failure 1: Duplicates

What happens: a job runs twice (a retry, a timeout, a double click, a webhook delivered twice) and creates two orders, two invoices or two customers.

Why: the receiving system cannot tell that it has already processed this event.

How to design against it:

  • Idempotency keys: give each operation a unique, stable identifier (such as the shop's order number). Before creating anything, check whether it already exists. Better, make the target reject duplicates.
  • Store the external ID of each record on both sides, so you always know what corresponds to what.
  • Use upserts (create or update) instead of blind creates.
  • Treat webhooks as at-least-once delivery: assume every event can arrive twice, or out of order.

For business readers, this is the reason you once got two invoices for one order. At higher volume, the same logic belongs in your own scripts rather than a no-code tool billed per run. That trade-off is covered in Automating without per-task pricing: scripts vs no-code.

Failure 2: Lost updates

What happens: a change in one system never reaches the other. A webhook was dropped while your server was down, an API returned an error and nobody retried, or a job crashed halfway.

How to design against it:

  • Retries with backoff for temporary errors
  • A queue or a log of pending changes, so nothing exists only in memory
  • Reconciliation jobs: a periodic comparison (nightly, for example) of both systems that finds and fixes differences. Webhooks give speed, reconciliation gives correctness. You want both.
  • Dead-letter handling: items that fail repeatedly go to a place where a human can see them, instead of vanishing
  • Alerts when the queue grows or a job has not run

Need a sync that survives retries?

We design shop and ERP integrations with source-of-truth rules, queues, and reconciliation, not a fragile one-shot webhook.

Failure 3: Conflicting updates

What happens: both systems change the same record close together. One overwrites the other, and someone's work disappears.

How to design against it:

  • One owner per field, as described above
  • If two-way sync is needed, use timestamps or version numbers and a clear rule: last write wins, source system wins, or conflicts go to a human
  • Avoid syncing derived fields (totals, statuses) that each system computes differently
  • Beware sync loops: an update in A triggers an update in B, which triggers an update in A. Break the loop by recording the origin of each change and ignoring changes you made yourself

Failure 4: Mismatched data models

What happens: the two systems do not describe the world the same way. One has "variants," the other has separate products. One stores a full name, the other first and last names. One has taxes included in prices, the other adds them. One counts "available" stock, the other counts "on hand."

How to design against it:

  • Write a mapping document field by field, with the transformation rules and examples
  • Handle the awkward cases explicitly: products with variants, bundles, partial refunds, discounts, multiple shipping addresses, tax rules by country
  • Normalise formats: dates and time zones, currencies and decimal handling (store money as integers or exact decimals, never floating point), phone and country formats, character encoding
  • Validate on the way in, and reject or flag records that do not fit rather than guessing
  • Test with real, ugly data: the customer with an apostrophe in their name, the order with 40 lines, the product without a SKU

Stock deserves special mention. "Available stock" depends on reservations, pending orders, returns and warehouses. Decide exactly what number you sync, and when.

Failure 5: Silent failure

What happens: the sync stops working, or works wrongly, and nobody notices for days or weeks. By the time someone sees it, there are hundreds of inconsistencies.

How to design against it:

  • Monitor the absence of success, not only errors: "no orders imported in 3 hours during business hours" should raise an alert
  • Track counts: records sent vs records received, per day
  • Log each run with a summary: processed, created, updated, skipped, failed
  • A visible dashboard or daily digest showing health at a glance
  • A manual re-sync tool for a single record, so support can fix cases without a developer
  • Expiring credentials: API tokens and OAuth connections expire. Track expiry dates and alert before they lapse

Webhooks, polling, or both?

  • Webhooks: the source tells you when something changes. Fast and efficient, but they can be lost or duplicated, and your endpoint must be available and secure (verify signatures).
  • Polling: you ask regularly for changes since your last check. Simpler to reason about, a little slower, and it uses API quota.
  • Both: webhooks for speed, a polling or reconciliation job as the safety net. This is the robust option for anything important.

Respect API rate limits: use batching where available, back off when you receive "too many requests," and avoid fetching everything on every run. Use "changed since" filters or cursors.

A minimal robust design

  • Events from the source go into a queue (or table) first, never processed inline in the webhook handler.
  • A worker processes events, idempotently, with retries.
  • Failures after N attempts go to a review list with alerts.
  • A nightly reconciliation compares both systems and repairs differences.
  • Every record stores its external ID and last sync time.
  • A dashboard or digest shows counts, errors and lag.
  • Secrets, tokens and keys stay in your vault, with expiry tracking.

Common pitfalls

  • Two-way sync everywhere "just in case"
  • No external IDs stored, so matching relies on names or emails
  • Trusting webhooks as the only mechanism
  • Processing the webhook inside the request, then timing out
  • Floating point for money
  • Time zone confusion causing off-by-one-day errors
  • No alerting on silence
  • Full re-sync every run, hitting rate limits
  • Hardcoding edge-case rules in several places

Checklist

  • Source of truth defined for each data type
  • Mapping document written, edge cases listed
  • External IDs stored on both sides
  • Operations are idempotent
  • Retries with backoff, and a review list for repeated failures
  • Reconciliation job in place
  • Conflict rules defined for any two-way field
  • Alerts on errors, silence and expiring credentials
  • Logs and a health summary available
  • Manual fix tool for single records

Related

From fragile sync to a design you can operate

If your shop and back office already disagree, or you are about to connect them, write us before the drift becomes a weekly firefight.