Beta

Create a tenant

A new tenant starts with its own users, OAuth settings, audit history and logs. You are its first Tenant Admin.

BTL Admin

Operating and monitoring an integration

The printer's photo connection has been running for a year. Most nights nothing happens that anyone notices. Then, early one December morning, the photo service slows down while forty thousand calendar jobs try to refresh at once, and the printer's on-call engineer has to decide quickly which failures to retry, which to report, and which customers really need to reconnect.

Errors and denied access handled a single failed connection: describe the outcome honestly and offer a fresh attempt. Running an integration at scale needs the same honesty written down as policy, before the night it is needed.

Deciding when to retry

A retry is safe when repeating the request cannot do harm. For OAuth, that depends on two questions. Did the request reach the server? And is anything it carries single-use?

SituationRetry automatically?What to do
The connection failed before the request was sentYes, with backoffNothing reached the server, so repeating is safe, even for a code exchange.
429 or 503 on a refresh or client credentials requestYes, after a delayThe server is asking for time. Wait at least as long as any Retry-After header says.
A code exchange was sent, then a timeout or a server errorNoThe code may already be spent, and it expires within minutes anyway. Offer a fresh attempt.
A refresh was sent, then a timeout or a 500At most onceWithout rotation, the same refresh token still works. With rotation, the retry succeeds only if the first request never took effect.
invalid_grant on a refreshNoMark the connection as needing reconnection.
invalid_clientNoAlert the operators. The deployed credential no longer matches the registration.
invalid_request, invalid_scope, unauthorized_clientNoA bug or a configuration change. Alert, and do not repeat the request.
server_error or temporarily_unavailable in an authorization responseNoTell the person and let them try again later, rather than redirecting them in a loop.

The refresh row needs care because the photo service rotates refresh tokens. If the first request did take effect, the retry presents a token the server has already replaced. Unless the server allows the immediately previous token for a few seconds, as Refresh tokens and rotation described, it treats the retry as reuse and revokes the family. That costs the printer nothing it could still use: with the replacement lost, the connection needed reconnecting from that moment. The event does land in the photo service's reuse alerts, so the printer tags such retries in its own logs, and both companies can tell a lost response from a stolen token.

Backoff should grow with each attempt, stop after a few tries, and include some randomness, called jitter, so that thousands of clients that failed together do not retry together. The calendar jobs make the point. If every one is scheduled for 03:00, every one meets the same slow token endpoint and then retries on the same schedule. Spreading the jobs across the night does more for reliability than any retry policy.

What to watch

The printer tracks a small set of signals, each counted rather than logged in full, and the photo service watches the matching ones on its side:

SignalWhat a change usually means
Token endpoint errors, by error code and clientA jump in invalid_client just after a deployment is a credential mismatch. A rise in invalid_grant on code exchanges suggests slow callbacks or a correlation bug.
Refresh failures, by reasonA steady trickle of invalid_grant is people disconnecting. A sudden rise can mean the photo service ended grants after a security event, or the printer is presenting stale tokens.
Refresh token reuse, as the photo service reports itA concurrency bug in the refresh coordinator, or refresh tokens in someone else's hands.
Credential and key expiry datesClient secrets with an expiry date, private keys, client certificates, and TLS certificates all expire on a schedule. Alert weeks ahead, not on the day.
Metadata and key set fetch failuresThe client or API is running on a cached copy and will miss the next change.
Token endpoint latency and error rateTrouble at the authorization server, often visible here before customers report it.

Each signal needs error codes, client IDs, and counts. None of them needs a token, and a dashboard that would only work with tokens in its data is a dashboard to redesign.

When the authorization server is down

When the photo service's authorization server is unavailable, some things keep working and others cannot. Access tokens the printer already holds remain valid until they expire, and an API that validates them locally keeps accepting them, although one that relies on introspection cannot. Nothing new can be obtained: no new connections and no refreshes.

The printer's behavior follows from that. A refresh that fails with a 503, or never reaches the server, leaves the connection as it is, connected but waiting, and the job tries again later with backoff. A refresh that reached the server and then timed out is a different case, because the server may already have rotated the token, so the refresh row of the retry table applies. Telling customers to reconnect would be wrong: their grants are fine, and reconnecting would fail too. The December calendar job reschedules itself instead of giving up. The Connect button still works, but leads to a clear message that the photo service is unavailable rather than to a broken redirect.

What the printer must not do is relax its own rules to stay up. Using access tokens past their expiry, or skipping the issuer check because metadata could not be refreshed, would turn an outage at someone else's service into a security problem at the printer's.

Incidents and records

Some incidents belong to the printer. If its client credential appears in a public repository, Rotating client credentials set out the order: revoke first, deploy the replacement, then review what was issued with the old credential. Operations adds the preparation. A written runbook names who can revoke the credential at the photo service at three in the morning, how to reach the photo service's security contact, and how to search the printer's logs by credential and time. Practicing it once, on the test client, finds the step that would otherwise fail during the real thing.

That review depends on logs written long before anything went wrong and kept for a deliberate period. Operational records of token requests, refreshes, and API failures are often kept for weeks or a few months, long enough to investigate an incident that is noticed late. Security events, such as credential changes and revocations, are usually kept longer. The right periods depend on each organization's obligations, but they should be chosen, written down, and enforced by automatic deletion rather than left to a logging service's defaults.

Keeping secrets out of logs is what makes that retention acceptable. A year of records containing codes, refresh tokens, or client secrets would be a store of credentials waiting to leak. A year of correlation IDs, client IDs, error codes, and token identifiers is an investigator's best resource, and on most nights it is simply a quiet record that the integration did its job.

Try it in the Lab

PUT IT INTO PRACTICE

Check your understanding

Try these questions before moving on. If an answer isn't right, use the feedback and try again.

0 of 2 answered correctly

Enable JavaScript to answer these questions and save progress in this browser.

QUESTION 1 OF 2The printer's code exchange times out after the request was sent. What should happen?

QUESTION 2 OF 2During a photo service outage, the calendar job's refresh requests return 503. What should the printer do?

We value your privacy

We use cookies and similar technologies to enhance your browsing experience, and analytics to understand our traffic. By clicking "Allow All", you consent to optional analytics. Cookie Policy

Learn identity