Discuss a project

Systems & integrations5 min read

API integration: what happens when the data does not match

An integration does not break when the connection drops. It breaks when two systems quietly show different things. This guide covers how duplicates appear, why retrying without an identifier is dangerous, and how to agree in advance which system counts as the truth.

Author
Devnora team
Published
Updated
Reading time
5 min read
Language
Read in Lithuanian
In this article
  1. In short
  2. Why a quiet failure is worse than a loud one
  3. How duplicates appear
  4. An idempotency key: the simplest protection
  5. Which system is the truth
  6. Reconciliation: noticing a mismatch yourself
  7. An illustrative example
  8. Boundaries of responsibility
  9. What this guide does not tell you

The worst state for an integration is not a dropped connection. A dropped connection announces itself: something stops working and somebody calls. The worst state is the quiet one — both systems are answering, nothing shows an error, and the numbers in them no longer agree. That kind of mismatch usually surfaces at month end, by which point nobody knows which figure is right or since when.

In short

  • Agree which system is the source of record for each kind of data. Without that agreement no mismatch can be resolved.
  • Retrying after a failure is necessary, but without an identifier it manufactures duplicates.
  • A message can be delivered twice — design so that the second time changes nothing.
  • Mismatches have to surface on their own rather than be explained after the fact. That needs at least a minimal reconciliation check.
  • The boundaries between systems are also boundaries of responsibility: write down who fixes what before the first incident.

Why a quiet failure is worse than a loud one

When an integration fails completely, the problem has an owner and a timestamp. When it works partially, the problem accumulates. The typical cases: a message was sent but the acknowledgement never arrived, so the sender assumes failure and repeats; a field was renamed in one system while the other keeps writing to the old one; two people edited the same record at the same time and whoever saved last won. None of these puts an error on anybody's screen.

How duplicates appear

Most integrations deliver at-least-once rather than exactly-once. In practice that means the same message can arrive a second time — after a network fault, after a retry, after a restart. If the receiving side treats every message it receives as a new fact, you get two records instead of one. This is not a theoretical risk; it is ordinary network behaviour.

  • The sender missed the acknowledgement and repeated, although the first attempt was recorded.
  • An intermediate service replayed the message on its own schedule.
  • A manual replay after an incident, without knowing how much had already been processed.
  • Two different sources reported the same event in different formats, and they were not recognised as the same.
  • An import was run a second time because the first looked like it had failed.

An idempotency key: the simplest protection

The fix is simple and often skipped: every message should carry a stable identifier created by the sender, and the receiving side should remember it. If the same identifier arrives again, the action is not repeated and the original response is returned. What matters is that the identifier belongs to the event rather than to the attempt — otherwise every retry looks like a new event.

  • The sender creates the identifier and it does not change across retries of the same event.
  • The receiver stores processed identifiers for as long as a retry can realistically be delayed.
  • A repeated identifier returns the same result rather than an error — otherwise the sender retries forever.
  • Retries use increasing intervals rather than firing immediately and indefinitely.
  • There is a limit after which retrying stops and the message is handed to a person — silently dropping it is not an option.

Which system is the truth

This agreement matters more than any technical detail. Until it is clear which system is the source of record for a given kind of data, every mismatch becomes a discussion. It helps to agree not about systems in general but about each kind of data separately: customer contact details might be owned in one system while order status is owned in another. Then a mismatch stops being a matter of opinion and becomes a correction.

  • For each kind of data: which system owns it and which merely reflects it.
  • What happens if it was changed in both places — the source of record wins, not whoever saved later.
  • Which fields are not synchronised at all, so nobody assumes they always agree.
  • What identity means for a record: which field decides that this is the same customer or order.
  • What happens when a record is deleted in the source of record.

Reconciliation: noticing a mismatch yourself

Even a well-designed integration will diverge eventually. The difference between a tidy and an untidy system is not whether that happens but whether it is noticed without a customer phoning. A minimal reconciliation is not much work: periodically compare record counts and a few critical fields across both sides, and have somewhere that a mismatch gets written down.

  • A regular count comparison: how many records exist on each side for the same period.
  • A comparison of a few critical fields rather than all of them — comparing everything usually never gets started.
  • A queue of unprocessed messages with a visible number, so a growing backlog is obvious.
  • Records of what was transferred and when — without them an incident cannot be investigated.
  • A clear alert to a named person when a mismatch exceeds an agreed threshold.

An illustrative example

Illustrative scenario, not a description of a client project. Imagine orders being passed from a website into accounting software. One night the network drops briefly after the accounting system has already recorded an order but before the website receives the acknowledgement. The website assumes failure and retries. Without an identifier, accounting now holds two identical orders and the customer receives two invoices. With one, the second attempt is recognised and the first result is returned. The point is that the problem was not the dropped connection — that is normal — but that the retry was not safe.

Boundaries of responsibility

A technical boundary between systems is almost always an organisational boundary too. Before launch it is worth writing down a few things that become the important ones during an incident: who watches the queue, who is allowed to replay a transfer, who may correct data directly, and who gets called if the mismatch originated on an external supplier's side. The list looks bureaucratic right up until the first time you need it.

What this guide does not tell you

It contains no figures and no recommendation of a specific technology: the same principles apply to queues, to direct interfaces and to file transfers. Nor does it give a universal reconciliation threshold — how much divergence counts as normal depends on the operation, and any number written here would be an assumption rather than a norm. Its purpose is different: that retries are safe, that the source of record is agreed, and that a mismatch is noticed without a customer phoning.

If you settle just two things before integration work begins — which system is the truth for each kind of data, and how a repeated message is recognised — most of the quiet failures have nowhere left to come from.

Share

Send by email

Next step

Related service

Business software development

Custom business systems and customer portals with clear workflows, access permissions, integrations and a phased approach to data migration.

If this article describes your situation, tell us what is not working. We will say whether and how we can help.

Worth reading next

  1. Systems & integrations

    Who owns an integration once it works?

    Integrations rarely break on day one. They break six months later, when the other system changes something and nobody had ever answered the question of who would notice. What to settle before the work rather than after.

More on this topic: Systems & integrations