Beta

Create a tenant

A new tenant starts with its own users, OAuth settings, audit history and logs. You are its first Tenant Admin.

BTL Admin

OAUTH 2.0 · LAB

Write a retry policy, test its rows, and read the signals

Turn the lesson's retry table into code, trigger its rows with real tenant answers and a labeled stand-in for an outage, rehearse a credential incident, and assemble the day's signals.

Partly readyIncludes a simulationUses your lab tenant

The lesson

Builds on: Implementing a client.

New to the labs? Start with the lab toolkit and the shared cast and names every lab uses.

Partly ready. Most of this lab runs today. Steps that wait on platform features are marked, and Missing infrastructure says what they need.

Your progress

Press Start before you begin. Only events your tenant records after that count, in the order below. Checking reads your tenant's Audit, so you need Audit read access in it.

  1. Retry a code exchange that was already sent

    Recorded as oauth.token rejected (code_replayed) for lab-printer.

  2. Retry a refresh after a lost response

    Recorded as oauth.token rejected (refresh_replayed) for lab-printer.

  3. Rotate the test client's secret for the incident rehearsal

    Recorded as tenant.oauth.credentials.rotate succeeded.

  4. See the undeployed secret refused

    Recorded as oauth.token rejected (invalid_client) for lab-tmp-printer-test.

Setup

  1. Choose Lab Photos as the lab tenant and press Start.

  2. Create lab-tmp-printer-test again as in Implementing a client Setup, read -rs PRINTER_TEST_SECRET; export PRINTER_TEST_SECRET, update client_id in ~/lab-printer-client/test.json, and connect Ava with node printer.mjs test.json connect ava. Work in ~/lab-printer-client.

  3. Load CLIENT_ID and CLIENT_SECRET for lab-printer with the helpers from Rotate a refresh token family. Confirm Default access tokens has a reuse grace of 0.

  4. Write the lesson's retry table down as code. Save it as retry.mjs:

// The printer's retry policy. sent: did the request reach the server? Every answer is one action.
export function decide({ kind, sent, status, error, attempt = 0 }) {
  if (!sent) return attempt < 4 ? 'retry_with_backoff' : 'reschedule';           // nothing reached the server
  if (kind === 'code_exchange') return error === 'invalid_client' ? 'alert_operators' : 'offer_fresh_attempt';   // a sent code may be spent
  if (status === 429 || status === 503) return attempt < 4 ? 'retry_after_delay' : 'reschedule';
  if (error === 'invalid_grant') return 'mark_needs_reconnection';
  if (['invalid_client', 'invalid_request', 'invalid_scope', 'unauthorized_client'].includes(error)) return 'alert_operators';
  if (kind === 'refresh' && (!status || status >= 500)) return attempt < 1 ? 'retry_once_tagged' : 'reschedule';
  return 'reschedule';
}
// Delays grow, stop at a minute, and carry jitter so clients that failed together do not retry together.
export const backoff = attempt => Math.round(Math.min(60_000, 1000 * 2 ** attempt) * (0.5 + Math.random() / 2));
  1. Add a small classifier that reads a curl result (body, then a status line; status 000 means nothing was sent):

classify() {  # $1 = kind, $2 = curl output ending in a status line
  local status body; status=$(tail -n1 <<<"$2"); body=$(sed '$d' <<<"$2")
  node --input-type=module -e "import { decide } from './retry.mjs'; console.log('$1', '$status', process.argv[1] || '-', '->', decide({ kind: '$1', sent: '$status' !== '000', status: Number('$status') || undefined, error: process.argv[1] || undefined }))" "$(jq -r '.error // empty' <<<"$body" 2>/dev/null)"
}

Walkthrough

  1. A code exchange that was already sent. Get a code for lab-printer (authorize "photos.read offline_access", sign in as Ava), redeem it with exchange, then send exactly the same request again, as a retry after a timeout would:

R=$(curl -s -w '\n%{http_code}' -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=authorization_code --data-urlencode "code=$CODE" \
  --data-urlencode "redirect_uri=http://127.0.0.1:8765/callback" -d "code_verifier=$VERIFIER"); classify code_exchange "$R"

invalid_grant, Audit reason code_replayed, and the policy answers offer_fresh_attempt. The tokens from the first exchange are now revoked too.

Why it matters: a retried code exchange can only make things worse. The person gets a fresh attempt instead.

  1. A refresh whose response was lost. Using the refresh token from step 1, refresh and throw the response away, then retry with the same token:

curl -s -o /dev/null -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=refresh_token --data-urlencode "refresh_token=$REFRESH"
R=$(curl -s -w '\n%{http_code}' -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=refresh_token --data-urlencode "refresh_token=$REFRESH"); classify refresh "$R"
echo "{\"at\":\"$(date -u +%FT%TZ)\",\"stage\":\"refresh_retry\",\"tag\":\"lost_response_retry\"}" >> retries.log

With a reuse grace of 0, the retry is refresh_replayed and the family ends; the policy answers mark_needs_reconnection. The tagged line in your own log lets both sides tell a lost response from a stolen token.

Why it matters: with the replacement lost, the connection needed reconnecting from that moment. A provider can soften this with a few seconds of reuse grace, which is a decision to make with them, not a retry loop to build.

  1. Rehearse a credential incident on the test client. The runbook's order is revoke, deploy, review:

    • Revoke: in Access Token Management, rotate lab-tmp-printer-test's secret, as you would if the old one appeared in a public repository. Do not update PRINTER_TEST_SECRET yet.

    • Observe: node printer.mjs test.json race ava is refused with 401 invalid_client; feed that to the policy with classify refresh "$(printf '{"error":"invalid_client"}\n401')": alert_operators, never a retry.

    • Deploy: read -rs PRINTER_TEST_SECRET with the new value, export it, and reconnect.

    • Review: in Audit, filter by client lab-tmp-printer-test and the last hour, and list what was issued with the old secret.

Note: the tenant has no overlap period for client secrets (G10), so the replacement must reach every running copy before the next request.

  1. Rows the tenant cannot be made to produce on request.

Simulation. the tenant cannot be made slow or unavailable for a lab, so a local stand-in answers 503 and a closed port stands in for a connection failure. The policy code under test is your own, unchanged.

node -e 'require("http").createServer((q,s)=>{s.writeHead(503,{"Retry-After":"30","Content-Type":"application/json"});s.end("{\"error\":\"temporarily_unavailable\"}")}).listen(8799,"127.0.0.1")' &
classify refresh "$(curl -s -w '\n%{http_code}' -X POST http://127.0.0.1:8799/oauth/token)"   # retry_after_delay
classify refresh "$(curl -s -w '\n%{http_code}' -X POST http://127.0.0.1:1/oauth/token)"      # retry_with_backoff
kill %1
node --input-type=module -e "import { backoff } from './retry.mjs'; console.log([0,1,2,3,4].map(backoff))"

During an outage the connection stays connected but waiting, the calendar job reschedules itself, and nobody is told to reconnect, because their grants are fine.

  1. Assemble the day's signals from real records. In Audit, filter Outcome: rejected for today. In Logs, open the protocol summaries. Count, with the first and latest request ID for each: invalid_client by client, invalid_grant by reason (code_replayed, refresh_replayed, refresh_revoked, pkce_failed), and rate_limited.

Note: there are no tenant metrics or alerts yet (G36), so these counts are assembled by hand. A jump in invalid_client right after step 3 is exactly what an alert should catch.

  1. Check retention. Audit and Logs keep 30 days. Note the Logs coverage line for any summarized or suppressed requests, and write down how long you would keep operational records and security events, and why.

Break it

Steps 1 to 4 are the failures, each restored in place: the secret is deployed in step 3 and the stand-in stopped in step 4. Nothing else in the tenant is weakened.

Check your work

Press Check my progress before Cleanup, because Cleanup deletes the test client. The checks look for, in order: code_replayed from the retried exchange, refresh_replayed from the lost-response retry, the secret rotation, and the refused request with the undeployed secret.

Your retries.log has one tagged retry, and no line in it contains a code or token.

Cleanup

  1. Make sure the stand-in on port 8799 is stopped.

  2. Delete retries.log and stored tokens: rm -f retries.log; rm -rf ~/lab-printer-oauth.

  3. Delete the client lab-tmp-printer-test. Run unset PRINTER_TEST_SECRET REFRESH CODE VERIFIER.

Missing infrastructure

  • G36 Tenant metrics and alerts. Counts over time by operation, reason and client, with thresholds that notify an administrator, such as a jump in invalid_client after a deployment or any refresh_replayed. Step 5 would then set an alert and see it fire after step 3.

  • G10 Client secret rotation overlap. With a short period where both the old and new secret work, step 3 becomes a zero-downtime rotation rehearsal: deploy the new secret, confirm it is in use, then retire the old one.

Back to all labs

We value your privacy

We use cookies and similar technologies to enhance your browsing experience, and analytics to understand our traffic. By clicking "Allow All", you consent to optional analytics. Cookie Policy

The Lab