Operating and monitoring an integration
The printer's photo connection has been running for a year. Most nights nothing happens that anyone notices. Then, early one December morning, the photo service slows down while forty thousand calendar jobs try to refresh at once, and the printer's on-call engineer has to decide quickly which failures to retry, which to report, and which customers really need to reconnect.
Errors and denied access handled a single failed connection: describe the outcome honestly and offer a fresh attempt. Running an integration at scale needs the same honesty written down as policy, before the night it is needed.
Deciding when to retry
A retry is safe when repeating the request cannot do harm. For OAuth, that depends on two questions. Did the request reach the server? And is anything it carries single-use?
| Situation | Retry automatically? | What to do |
|---|---|---|
| The connection failed before the request was sent | Yes, with backoff | Nothing reached the server, so repeating is safe, even for a code exchange. |
| 429 or 503 on a refresh or client credentials request | Yes, after a delay | The server is asking for time. Wait at least as long as any Retry-After header says. |
| A code exchange was sent, then a timeout or a server error | No | The code may already be spent, and it expires within minutes anyway. Offer a fresh attempt. |
| A refresh was sent, then a timeout or a 500 | At most once | Without rotation, the same refresh token still works. With rotation, the retry succeeds only if the first request never took effect. |
invalid_grant on a refresh | No | Mark the connection as needing reconnection. |
invalid_client | No | Alert the operators. The deployed credential no longer matches the registration. |
invalid_request, invalid_scope, unauthorized_client | No | A bug or a configuration change. Alert, and do not repeat the request. |
server_error or temporarily_unavailable in an authorization response | No | Tell the person and let them try again later, rather than redirecting them in a loop. |
The refresh row needs care because the photo service rotates refresh tokens. If the first request did take effect, the retry presents a token the server has already replaced. Unless the server allows the immediately previous token for a few seconds, as Refresh tokens and rotation described, it treats the retry as reuse and revokes the family. That costs the printer nothing it could still use: with the replacement lost, the connection needed reconnecting from that moment. The event does land in the photo service's reuse alerts, so the printer tags such retries in its own logs, and both companies can tell a lost response from a stolen token.
Backoff should grow with each attempt, stop after a few tries, and include some randomness, called jitter, so that thousands of clients that failed together do not retry together. The calendar jobs make the point. If every one is scheduled for 03:00, every one meets the same slow token endpoint and then retries on the same schedule. Spreading the jobs across the night does more for reliability than any retry policy.
What to watch
The printer tracks a small set of signals, each counted rather than logged in full, and the photo service watches the matching ones on its side:
| Signal | What a change usually means |
|---|---|
| Token endpoint errors, by error code and client | A jump in invalid_client just after a deployment is a credential mismatch. A rise in invalid_grant on code exchanges suggests slow callbacks or a correlation bug. |
| Refresh failures, by reason | A steady trickle of invalid_grant is people disconnecting. A sudden rise can mean the photo service ended grants after a security event, or the printer is presenting stale tokens. |
| Refresh token reuse, as the photo service reports it | A concurrency bug in the refresh coordinator, or refresh tokens in someone else's hands. |
| Credential and key expiry dates | Client secrets with an expiry date, private keys, client certificates, and TLS certificates all expire on a schedule. Alert weeks ahead, not on the day. |
| Metadata and key set fetch failures | The client or API is running on a cached copy and will miss the next change. |
| Token endpoint latency and error rate | Trouble at the authorization server, often visible here before customers report it. |
Each signal needs error codes, client IDs, and counts. None of them needs a token, and a dashboard that would only work with tokens in its data is a dashboard to redesign.
Incidents and records
Some incidents belong to the printer. If its client credential appears in a public repository, Rotating client credentials set out the order: revoke first, deploy the replacement, then review what was issued with the old credential. Operations adds the preparation. A written runbook names who can revoke the credential at the photo service at three in the morning, how to reach the photo service's security contact, and how to search the printer's logs by credential and time. Practicing it once, on the test client, finds the step that would otherwise fail during the real thing.
That review depends on logs written long before anything went wrong and kept for a deliberate period. Operational records of token requests, refreshes, and API failures are often kept for weeks or a few months, long enough to investigate an incident that is noticed late. Security events, such as credential changes and revocations, are usually kept longer. The right periods depend on each organization's obligations, but they should be chosen, written down, and enforced by automatic deletion rather than left to a logging service's defaults.
Keeping secrets out of logs is what makes that retention acceptable. A year of records containing codes, refresh tokens, or client secrets would be a store of credentials waiting to leak. A year of correlation IDs, client IDs, error codes, and token identifiers is an investigator's best resource, and on most nights it is simply a quiet record that the integration did its job.