API integrations: why connecting the services is not enough
An integration has to be observable: events, statuses, logs, failures and recovery.
An integration is usually considered done once data has travelled from one system to another. In a demo that looks convincing: the request appeared in the CRM, the status updated, everyone is happy.
The problems start later and look different. Not all of the data arrived. Something came twice. An external service returned an error nobody heard about. A week later a discrepancy surfaces, and reconstructing what happened is impossible — there are neither traces nor statuses.
What an event is
Designing an integration starts not with endpoints but with a list of business events: a client sent a request, a payment went through, a manager changed a status, a document was signed. An event is something that happened in the business, not a method call.
That list is useful before any code. It shows which systems need to hear about an event at all, what happens when one of them is unavailable, and which events can be lost without consequence — and which cannot be lost under any circumstances.
For each event it pays to agree on the payload up front: what is always sent, what may be absent, what must never change. Format changes six months later are normal, and only the integration where the event version is explicit — and the receiving side ignores unknown fields — will survive them.
Why statuses are needed
Transferring data is not an instant action but a process with states. A message is accepted, queued, sent, confirmed, rejected. Until those states exist, the only available answer to “did it arrive” is “probably”.
- every transfer has an identifier of its own
- the current state and the time of the last attempt are visible
- resending the same event does not create a duplicate
- there are terminal states, not only “in progress”
How to handle failures
Failures in integrations fall into two unequal groups. Temporary ones — the network, a timeout, an overloaded service on the other side: a retry with a growing pause cures those. Permanent ones — a wrong format, a missing field, a rule-based rejection: retrying them is pointless however many times you try.
They have to be separated at the code level, otherwise the system spends hours battering a wall while staying silent about the real problem. Permanent failures should land where a person will see them, and stay in a state they can be restarted from once fixed.
Retries require one condition that is often forgotten: the receiving side has to tolerate the same event arriving twice. Otherwise the first failed attempt that actually got through creates a duplicate request or a second payment. Testing it is simple — send one event twice and look at what happened.
It pays to decide in advance what happens to messages that could not be delivered at all. Losing them is not an option; handling them automatically is not either. They usually go into a separate queue where they wait for a person: they look at the reason, fix the data and restart it. Without such a queue, failed messages either vanish or get stuck retrying forever.
Logs and documentation
An integration log is written not for the developer but for whoever will investigate six months later. A useful entry answers: which event, where it was sent, what came back, when, and which attempt this was. Personal data does not belong in logs — an identifier is enough.
Documentation is needed in roughly the same volume: the list of events, the format of each, the retry rules, what counts as a failure and who to go to when it breaks. One page written straight away saves weeks a year later, when the team has changed.
It is worth writing down separately what to do about the usual failures: how to restart a stuck transfer, how to find a particular request by its identifier, who to call when the external service has been erroring for three hours. That is an instruction for whoever is on duty, not for the developer, and it has to be written so that someone who did not build the integration can follow it.
Supporting the integration
An integration is never finished for good. On the other side API versions change, new required fields appear, rules shift. So it needs an owner and a way to learn about a breakage before the client reports it.
The minimum is an alert when errors pile up and a simple screen showing the queue and the latest failures. That is enough to find a discrepancy within an hour rather than a week later through a complaint.
A regular reconciliation helps too: once a day compare the number of records on both sides and surface the difference. It catches what no log will — events that were never sent at all because nobody created them. Those losses are the nastiest: from the integration’s side everything looks perfect.
Systems connected, but failures found too late?
Describe which services exchange data and what happens when something fails. We will work through the events, statuses, retries and what is missing for observability.
Show us your process