Beta

Create a tenant

A new tenant starts with its own users, OAuth settings, audit history and logs. You are its first Tenant Admin.

BTL Admin

Keeping systems in sync

A request with no answer

Go back to Monday 26 October, when the provisioning service created Dana Okafor's patient records account. Suppose that this time the request had gone differently. At 09:14 the provisioning service sends POST /Users with Dana's details. Thirty seconds pass without a response, and the connection times out.

Does Dana have an account? The provisioning service cannot tell. The request may never have reached the records system. It may have arrived and failed. Or the records system may have created the account and sent back 201 Created, and the response was lost on the way. From the client's side, all three look the same: silence.

Both obvious reactions are wrong. If the client treats the timeout as a failure and sends the same create again, and the first one did succeed, the second meets the account the first created. Because userName must be unique, it comes back with 409 Conflict, and a client that answers a conflict by inventing a new name creates dana.okafor2, a second account for the same nurse. If instead the client treats the silence as success, it has no id to store, and if the request never arrived, Dana starts work with no records account at all.

The accurate description is that the outcome is unresolved. The change might or might not have happened, and the client has to find out which before it does anything else.

Safe retries

Some requests can be repeated without harm. Setting Dr. Lee Moreau's active to false twice leaves the account disabled, and adding Sam Reyes to Oncology nurses twice leaves Sam a member once. An operation that has the same effect however many times it runs is idempotent, and a client can simply send it again after a timeout. Removing a member has the same effect when repeated, but the reply may differ: as SCIM requests and responses showed, some services answer a second removal with 400 and noTarget, which the client treats as done once it confirms the member is gone. A create is not idempotent, because each successful create makes another account.

Some APIs let a client attach a unique key to a create, so that the server recognizes a repeat and returns the original result. SCIM defines no such key, so the client has to check for itself, and externalId is what makes that possible. The identity system put E10482 in the create, so it can ask the records system whether a user with that key exists:

GET /scim/v2/Users?filter=externalId%20eq%20%22E10482%22 HTTP/1.1
Host: records.harborclinic.example
Authorization: Bearer demo-provisioning-token-7

HTTP/1.1 200 OK
Content-Type: application/scim+json

{
  "schemas": ["urn:ietf:params:scim:api:messages:2.0:ListResponse"],
  "totalResults": 1,
  "startIndex": 1,
  "itemsPerPage": 1,
  "Resources": [
    {
      "id": "a3f1c9e2-58d4-4b7e-9c61-2e7d0b4f8a15",
      "externalId": "E10482",
      "userName": "dana.okafor",
      "active": false
    }
  ]
}

All names, hosts, identifiers, and tokens in these examples are fictional.

Here the search finds the account, so the create worked and only the response was lost. The client stores the returned id, compares the account with what it meant to create, and sends any differences as an update. If the search had found nothing, the create never happened and sending it again would be safe. If the search times out too, the outcome stays unresolved and the client tries the search again later. Dana's other changes wait in the meantime, since there is no point adding Dana to a group before the client knows the account exists.

The provisioning service sends POST /Users for Dana to the patient records SCIM service. No response arrives before the timeout, so the outcome is unresolved rather than failed. The client searches with GET /Users and the filter externalId eq E10482. If the search finds the user, the client stores the returned id and does not send a second create. If the search finds nothing, the client sends the create again and receives 201 Created. The provisioning service sends POST /Users for Dana to the patient records SCIM service. No response arrives before the timeout, so the outcome is unresolved rather than failed. The client searches with GET /Users and the filter externalId eq E10482. If the search finds the user, the client stores the returned id and does not send a second create. If the search finds nothing, the client sends the create again and receives 201 Created.
After a timeout, the client looks for its own key before doing anything else. Only a search that finds nothing makes a second create safe.

One gap remains. A slow first request could still be inside the records system while the search runs, and finish just after the search reports nothing. Waiting a little longer than the records system's own processing time before searching makes that unlikely, and the unique userName makes it harmless: only one of the two creates can claim dana.okafor, and the other gets 409. The client settles that conflict by searching for its externalId again, never by choosing a new name.

Throughout, the provisioning service keeps a record of each intended change under an ID of its own, such as chg-20261026-0412 for creating the records account of identity hc-7f3k2q. Every attempt goes into that record: when it was sent, what came back, the timeout, the search, and the final result. If the records system accepts a request ID header and writes it to its own logs, the provisioning service sends the change ID there as well, so the two sides' records can be matched later. That is an agreement between the two systems, not part of SCIM. Afterwards, anyone can see that Dana's account was unresolved for a few minutes and then confirmed, rather than reported as failed or quietly assumed to exist.

Order and timing

Retries solve one problem and can create another. A change that waits in a queue, or is retried after a delay, can arrive after a later change to the same account.

Suppose the medical staffing office had updated Dr. Moreau's department in the contractor register at 17:50 on Friday 27 November, ten minutes before the contract ended, and that the provisioning service sent each change as a whole user with PUT, which Harbor's does not. It builds a PUT from its copy of Dr. Moreau's records account, which still says "active": true, but a slow connection holds the request up. At 18:02 the deactivation goes out on another connection and succeeds. At 18:04 the delayed PUT finally arrives. Applied as sent, it would change the department and also set active back to true, enabling a leaver's account again.

A late update to Dr. Moreau's records account
TimeWhat happensAccount version
17:50The provisioning service builds a PUT with If-Match: W/"5". A slow connection holds it up.W/"5"
18:02The deactivation succeeds on another connection.W/"6"
18:04The delayed PUT arrives. Its version no longer matches, so the records system answers 412 Precondition Failed.W/"6"
18:04The client reads the account again and rebuilds its change from the current intended state. The contract has ended, so nothing is sent.W/"6"

The version check is what stops the late request. When the provisioning service built the PUT, it attached the version it had read. The deactivation changed the version, so the records system refused the stale request instead of applying it. As SCIM requests and responses described, the client then reads the account again. This time it does not simply reapply its old change. It asks the identity system what the account should look like now, and the answer is a disabled account for someone whose contract has ended.

Not every service supports ETags, and its /ServiceProviderConfig says whether it does. Without them, the client has to keep order itself. Two habits help either way. The client sends one change at a time for each account, so a later change waits until an earlier one has finished. And it builds each request from the current intended state when it sends it, not from a copy taken when the change was queued. A PATCH that touched only the department would also have been safer than a PUT carrying a stale copy of everything.

Reconciliation

Careful retries keep the provisioning service's own changes correct. They cannot see changes it did not make. Over weeks, the records system moves away from what the identity system intends, for three common reasons:

  • Manual edits. An administrator changes the application directly, perhaps to help someone in a hurry.
  • Failed calls. A request was rejected or never resolved, and nobody followed it up.
  • Changes made inside the target. The application changes accounts through its own rules, or its own administrators create accounts the identity system never hears about.

The gap between intended and actual state is drift. Reconciliation finds it by comparing the two and deciding what to do with each difference. Each night, the identity system reads every user and group from the records system, page by page, and matches each account to an identity by its stored id or its externalId. It never matches on a name or email address alone, which is how duplicates and wrong matches begin. Then it compares the attributes and memberships it is responsible for.

Part of one night's reconciliation report for the patient records system
AccountIntendedActualLikely causeAction
Sam Reyes, E07731Member of Oncology nursesAlso a member of Pediatrics nursesAdded back by hand in the records system after the handover endedRemove, and report to Morgan Hale
A ward clerk, E09356Department: Day surgeryDepartment: SurgeryThe move update was rejected because the records system did not yet know the new unit, and was never sent againFix: send the current value
A part-time nurse, E08812ActiveDisabledThe records system disabled the account after six months without a sign-inReport to Morgan Hale. Do not enable it again automatically
er.locum3, no externalIdNo matching identityActive, last sign-in 14 AugustCreated by hand inside the records systemTreat as an orphan

Each row needs a different response. The department is an attribute the identity system owns, so it fixes the difference by sending the current value, the same request it would have sent the first time. Sam's extra membership is access, and access added outside the normal process needs a person's attention as well as a fix. The identity system removes it and tells Morgan Hale, who owns access to patient records, who added it and when. If Sam really needs pediatric records again, that need goes through a request with a reason and an end date.

The part-time nurse's account shows that not every difference is an error to overwrite. The records system disabled it on purpose, under its own rule for unused accounts, and enabling it again automatically would undo a deliberate control. The difference is reported so that Morgan Hale can decide whether the nurse still needs the account.

The last row has no identity behind it. The report cannot say whose account er.locum3 is, whether a locum still uses it, or whether something depends on it, so disabling it blindly could break something. It is an orphaned account, and Orphaned, dormant, and shared accounts covers how to find its owner and decide what happens to it.

One kind of difference does not wait for a nightly report. An account that is active in the records system for someone the identity system says has left is access that should already be gone, so it raises an alert the moment reconciliation finds it. Reconciliation also covers applications the provisioning service never calls. Comparing the billing system's user list with the identity system each month catches tickets that were closed without the work being done.

Limits and backlogs

Applications go offline. On a Tuesday night in December, the records system is down for a four-hour upgrade. During those hours the clinic merges two units, which changes the department and group memberships of dozens of staff, and three people reach the end of their last shift. By the time the records system comes back, several hundred changes are waiting.

If the provisioning service sends them all at once, the records system protects itself by refusing some of them:

HTTP/1.1 429 Too Many Requests
Retry-After: 30

Neither 429 Too Many Requests nor the Retry-After header is specific to SCIM. Both come from HTTP itself, and any API can use them. The client waits at least as long as Retry-After says, here 30 seconds. When a refusal gives no wait time, the client uses backoff: it waits a little after the first refusal, longer after the second, and longer again after that, up to a limit, adding a small random delay so that many waiting requests do not all retry at the same moment.

A client that cannot send everything at once has to choose what goes first. Harbor's provisioning service sends deactivations for leavers first, then other removals such as access a mover no longer needs, then new accounts for joiners who start soon, and finally attribute updates such as the merged unit's name. A delayed unit name is an inconvenience. A delayed removal is a former member of staff who can still open patient records. Within one account, changes still go in order, and because each request is built from the current intended state, a leaver's queued updates collapse into a single deactivation.

Priority is not a guarantee. If a removal is still unresolved after a set time, the provisioning service raises an alert to Jordan Ellis and Morgan Hale. Jordan then disables the account directly in the records system and records doing so, and the next reconciliation confirms that the actual state matches the intended one. Disabling the person at the sign-in service first, as Leaving an organization described, stops new sign-ins in the meantime, though not a session that is already open.

Retries, version checks, reconciliation, and priorities all start from the identity system's intended state: what each person should have. Continue to Access rules and birthright access to see how that intended state is decided in the first place.

Try it in the Lab

PUT IT INTO PRACTICE

Check your understanding

Try these questions before moving on. If an answer isn't right, use the feedback and try again.

0 of 2 answered correctly

Enable JavaScript to answer these questions and save progress in this browser.

QUESTION 1 OF 2The provisioning service's POST /Users for Dana times out with no response. What should it do before sending the create again?

QUESTION 2 OF 2After an outage, several hundred provisioning changes are waiting for the records system, which is now answering some requests with 429. Three of the changes deactivate leavers. Which should go first?

We value your privacy

We use cookies and similar technologies to enhance your browsing experience, and analytics to understand our traffic. By clicking "Allow All", you consent to optional analytics. Cookie Policy

Learn identity