Keeping systems in sync
A request with no answer
Go back to Monday 26 October, when the provisioning service created Dana Okafor's patient records account. Suppose that this time the request had gone differently. At 09:14 the provisioning service sends POST /Users with Dana's details. Thirty seconds pass without a response, and the connection times out.
Does Dana have an account? The provisioning service cannot tell. The request may never have reached the records system. It may have arrived and failed. Or the records system may have created the account and sent back 201 Created, and the response was lost on the way. From the client's side, all three look the same: silence.
Both obvious reactions are wrong. If the client treats the timeout as a failure and sends the same create again, and the first one did succeed, the second meets the account the first created. Because userName must be unique, it comes back with 409 Conflict, and a client that answers a conflict by inventing a new name creates dana.okafor2, a second account for the same nurse. If instead the client treats the silence as success, it has no id to store, and if the request never arrived, Dana starts work with no records account at all.
The accurate description is that the outcome is unresolved. The change might or might not have happened, and the client has to find out which before it does anything else.
Safe retries
Some requests can be repeated without harm. Setting Dr. Lee Moreau's active to false twice leaves the account disabled, and adding Sam Reyes to Oncology nurses twice leaves Sam a member once. An operation that has the same effect however many times it runs is idempotent, and a client can simply send it again after a timeout. Removing a member has the same effect when repeated, but the reply may differ: as SCIM requests and responses showed, some services answer a second removal with 400 and noTarget, which the client treats as done once it confirms the member is gone. A create is not idempotent, because each successful create makes another account.
Some APIs let a client attach a unique key to a create, so that the server recognizes a repeat and returns the original result. SCIM defines no such key, so the client has to check for itself, and externalId is what makes that possible. The identity system put E10482 in the create, so it can ask the records system whether a user with that key exists:
GET /scim/v2/Users?filter=externalId%20eq%20%22E10482%22 HTTP/1.1
Host: records.harborclinic.example
Authorization: Bearer demo-provisioning-token-7
HTTP/1.1 200 OK
Content-Type: application/scim+json
{
"schemas": ["urn:ietf:params:scim:api:messages:2.0:ListResponse"],
"totalResults": 1,
"startIndex": 1,
"itemsPerPage": 1,
"Resources": [
{
"id": "a3f1c9e2-58d4-4b7e-9c61-2e7d0b4f8a15",
"externalId": "E10482",
"userName": "dana.okafor",
"active": false
}
]
}
All names, hosts, identifiers, and tokens in these examples are fictional.
Here the search finds the account, so the create worked and only the response was lost. The client stores the returned id, compares the account with what it meant to create, and sends any differences as an update. If the search had found nothing, the create never happened and sending it again would be safe. If the search times out too, the outcome stays unresolved and the client tries the search again later. Dana's other changes wait in the meantime, since there is no point adding Dana to a group before the client knows the account exists.
One gap remains. A slow first request could still be inside the records system while the search runs, and finish just after the search reports nothing. Waiting a little longer than the records system's own processing time before searching makes that unlikely, and the unique userName makes it harmless: only one of the two creates can claim dana.okafor, and the other gets 409. The client settles that conflict by searching for its externalId again, never by choosing a new name.
Throughout, the provisioning service keeps a record of each intended change under an ID of its own, such as chg-20261026-0412 for creating the records account of identity hc-7f3k2q. Every attempt goes into that record: when it was sent, what came back, the timeout, the search, and the final result. If the records system accepts a request ID header and writes it to its own logs, the provisioning service sends the change ID there as well, so the two sides' records can be matched later. That is an agreement between the two systems, not part of SCIM. Afterwards, anyone can see that Dana's account was unresolved for a few minutes and then confirmed, rather than reported as failed or quietly assumed to exist.
Order and timing
Retries solve one problem and can create another. A change that waits in a queue, or is retried after a delay, can arrive after a later change to the same account.
Suppose the medical staffing office had updated Dr. Moreau's department in the contractor register at 17:50 on Friday 27 November, ten minutes before the contract ended, and that the provisioning service sent each change as a whole user with PUT, which Harbor's does not. It builds a PUT from its copy of Dr. Moreau's records account, which still says "active": true, but a slow connection holds the request up. At 18:02 the deactivation goes out on another connection and succeeds. At 18:04 the delayed PUT finally arrives. Applied as sent, it would change the department and also set active back to true, enabling a leaver's account again.
| Time | What happens | Account version |
|---|---|---|
| 17:50 | The provisioning service builds a PUT with If-Match: W/"5". A slow connection holds it up. | W/"5" |
| 18:02 | The deactivation succeeds on another connection. | W/"6" |
| 18:04 | The delayed PUT arrives. Its version no longer matches, so the records system answers 412 Precondition Failed. | W/"6" |
| 18:04 | The client reads the account again and rebuilds its change from the current intended state. The contract has ended, so nothing is sent. | W/"6" |
The version check is what stops the late request. When the provisioning service built the PUT, it attached the version it had read. The deactivation changed the version, so the records system refused the stale request instead of applying it. As SCIM requests and responses described, the client then reads the account again. This time it does not simply reapply its old change. It asks the identity system what the account should look like now, and the answer is a disabled account for someone whose contract has ended.
Not every service supports ETags, and its /ServiceProviderConfig says whether it does. Without them, the client has to keep order itself. Two habits help either way. The client sends one change at a time for each account, so a later change waits until an earlier one has finished. And it builds each request from the current intended state when it sends it, not from a copy taken when the change was queued. A PATCH that touched only the department would also have been safer than a PUT carrying a stale copy of everything.
Reconciliation
Careful retries keep the provisioning service's own changes correct. They cannot see changes it did not make. Over weeks, the records system moves away from what the identity system intends, for three common reasons:
- Manual edits. An administrator changes the application directly, perhaps to help someone in a hurry.
- Failed calls. A request was rejected or never resolved, and nobody followed it up.
- Changes made inside the target. The application changes accounts through its own rules, or its own administrators create accounts the identity system never hears about.
The gap between intended and actual state is drift. Reconciliation finds it by comparing the two and deciding what to do with each difference. Each night, the identity system reads every user and group from the records system, page by page, and matches each account to an identity by its stored id or its externalId. It never matches on a name or email address alone, which is how duplicates and wrong matches begin. Then it compares the attributes and memberships it is responsible for.
| Account | Intended | Actual | Likely cause | Action |
|---|---|---|---|---|
Sam Reyes, E07731 | Member of Oncology nurses | Also a member of Pediatrics nurses | Added back by hand in the records system after the handover ended | Remove, and report to Morgan Hale |
A ward clerk, E09356 | Department: Day surgery | Department: Surgery | The move update was rejected because the records system did not yet know the new unit, and was never sent again | Fix: send the current value |
A part-time nurse, E08812 | Active | Disabled | The records system disabled the account after six months without a sign-in | Report to Morgan Hale. Do not enable it again automatically |
er.locum3, no externalId | No matching identity | Active, last sign-in 14 August | Created by hand inside the records system | Treat as an orphan |
Each row needs a different response. The department is an attribute the identity system owns, so it fixes the difference by sending the current value, the same request it would have sent the first time. Sam's extra membership is access, and access added outside the normal process needs a person's attention as well as a fix. The identity system removes it and tells Morgan Hale, who owns access to patient records, who added it and when. If Sam really needs pediatric records again, that need goes through a request with a reason and an end date.
The part-time nurse's account shows that not every difference is an error to overwrite. The records system disabled it on purpose, under its own rule for unused accounts, and enabling it again automatically would undo a deliberate control. The difference is reported so that Morgan Hale can decide whether the nurse still needs the account.
The last row has no identity behind it. The report cannot say whose account er.locum3 is, whether a locum still uses it, or whether something depends on it, so disabling it blindly could break something. It is an orphaned account, and Orphaned, dormant, and shared accounts covers how to find its owner and decide what happens to it.
One kind of difference does not wait for a nightly report. An account that is active in the records system for someone the identity system says has left is access that should already be gone, so it raises an alert the moment reconciliation finds it. Reconciliation also covers applications the provisioning service never calls. Comparing the billing system's user list with the identity system each month catches tickets that were closed without the work being done.
Limits and backlogs
Applications go offline. On a Tuesday night in December, the records system is down for a four-hour upgrade. During those hours the clinic merges two units, which changes the department and group memberships of dozens of staff, and three people reach the end of their last shift. By the time the records system comes back, several hundred changes are waiting.
If the provisioning service sends them all at once, the records system protects itself by refusing some of them:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
Neither 429 Too Many Requests nor the Retry-After header is specific to SCIM. Both come from HTTP itself, and any API can use them. The client waits at least as long as Retry-After says, here 30 seconds. When a refusal gives no wait time, the client uses backoff: it waits a little after the first refusal, longer after the second, and longer again after that, up to a limit, adding a small random delay so that many waiting requests do not all retry at the same moment.
A client that cannot send everything at once has to choose what goes first. Harbor's provisioning service sends deactivations for leavers first, then other removals such as access a mover no longer needs, then new accounts for joiners who start soon, and finally attribute updates such as the merged unit's name. A delayed unit name is an inconvenience. A delayed removal is a former member of staff who can still open patient records. Within one account, changes still go in order, and because each request is built from the current intended state, a leaver's queued updates collapse into a single deactivation.
Priority is not a guarantee. If a removal is still unresolved after a set time, the provisioning service raises an alert to Jordan Ellis and Morgan Hale. Jordan then disables the account directly in the records system and records doing so, and the next reconciliation confirms that the actual state matches the intended one. Disabling the person at the sign-in service first, as Leaving an organization described, stops new sign-ins in the meantime, though not a session that is already open.
Retries, version checks, reconciliation, and priorities all start from the identity system's intended state: what each person should have. Continue to Access rules and birthright access to see how that intended state is decided in the first place.