OAUTH 2.0 · LAB
Write a retry policy, test its rows, and read the signals
Turn the lesson's retry table into code, trigger its rows with real tenant answers and a labeled stand-in for an outage, rehearse a credential incident, and assemble the day's signals.
Partly readyIncludes a simulationUses your lab tenant
The lesson
Builds on: Implementing a client.
New to the labs? Start with the lab toolkit and the shared cast and names every lab uses.
Partly ready. Most of this lab runs today. Steps that wait on platform features are marked, and Missing infrastructure says what they need.
- G36 Detection, risk signals, tenant metrics and alerts (baselines, anomaly rules, thresholds that notify an admin)
- G10 Client secret rotation overlap
Your progress
Press Start before you begin. Only events your tenant records after that count, in the order below. Checking reads your tenant's Audit, so you need Audit read access in it.
Sign in to start this lab and check your progress. Log in or create an account.
Retry a code exchange that was already sent
Recorded as
oauth.tokenrejected (code_replayed) forlab-printer.Retry a refresh after a lost response
Recorded as
oauth.tokenrejected (refresh_replayed) forlab-printer.Rotate the test client's secret for the incident rehearsal
Recorded as
tenant.oauth.credentials.rotatesucceeded.See the undeployed secret refused
Recorded as
oauth.tokenrejected (invalid_client) forlab-tmp-printer-test.
Setup
Choose Lab Photos as the lab tenant and press Start.
Create
lab-tmp-printer-testagain as in Implementing a client Setup,read -rs PRINTER_TEST_SECRET; export PRINTER_TEST_SECRET, updateclient_idin~/lab-printer-client/test.json, and connect Ava withnode printer.mjs test.json connect ava. Work in~/lab-printer-client.Load
CLIENT_IDandCLIENT_SECRETforlab-printerwith the helpers from Rotate a refresh token family. Confirm Default access tokens has a reuse grace of0.Write the lesson's retry table down as code. Save it as
retry.mjs:
// The printer's retry policy. sent: did the request reach the server? Every answer is one action.
export function decide({ kind, sent, status, error, attempt = 0 }) {
if (!sent) return attempt < 4 ? 'retry_with_backoff' : 'reschedule'; // nothing reached the server
if (kind === 'code_exchange') return error === 'invalid_client' ? 'alert_operators' : 'offer_fresh_attempt'; // a sent code may be spent
if (status === 429 || status === 503) return attempt < 4 ? 'retry_after_delay' : 'reschedule';
if (error === 'invalid_grant') return 'mark_needs_reconnection';
if (['invalid_client', 'invalid_request', 'invalid_scope', 'unauthorized_client'].includes(error)) return 'alert_operators';
if (kind === 'refresh' && (!status || status >= 500)) return attempt < 1 ? 'retry_once_tagged' : 'reschedule';
return 'reschedule';
}
// Delays grow, stop at a minute, and carry jitter so clients that failed together do not retry together.
export const backoff = attempt => Math.round(Math.min(60_000, 1000 * 2 ** attempt) * (0.5 + Math.random() / 2));
Add a small classifier that reads a curl result (body, then a status line; status
000means nothing was sent):
classify() { # $1 = kind, $2 = curl output ending in a status line
local status body; status=$(tail -n1 <<<"$2"); body=$(sed '$d' <<<"$2")
node --input-type=module -e "import { decide } from './retry.mjs'; console.log('$1', '$status', process.argv[1] || '-', '->', decide({ kind: '$1', sent: '$status' !== '000', status: Number('$status') || undefined, error: process.argv[1] || undefined }))" "$(jq -r '.error // empty' <<<"$body" 2>/dev/null)"
}
Walkthrough
A code exchange that was already sent. Get a code for
lab-printer(authorize "photos.read offline_access", sign in as Ava), redeem it withexchange, then send exactly the same request again, as a retry after a timeout would:
R=$(curl -s -w '\n%{http_code}' -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=authorization_code --data-urlencode "code=$CODE" \
--data-urlencode "redirect_uri=http://127.0.0.1:8765/callback" -d "code_verifier=$VERIFIER"); classify code_exchange "$R"
invalid_grant, Audit reason code_replayed, and the policy answers offer_fresh_attempt. The tokens from the first exchange are now revoked too.
Why it matters: a retried code exchange can only make things worse. The person gets a fresh attempt instead.
A refresh whose response was lost. Using the refresh token from step 1, refresh and throw the response away, then retry with the same token:
curl -s -o /dev/null -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=refresh_token --data-urlencode "refresh_token=$REFRESH"
R=$(curl -s -w '\n%{http_code}' -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/token" -d grant_type=refresh_token --data-urlencode "refresh_token=$REFRESH"); classify refresh "$R"
echo "{\"at\":\"$(date -u +%FT%TZ)\",\"stage\":\"refresh_retry\",\"tag\":\"lost_response_retry\"}" >> retries.log
With a reuse grace of 0, the retry is refresh_replayed and the family ends; the policy answers mark_needs_reconnection. The tagged line in your own log lets both sides tell a lost response from a stolen token.
Why it matters: with the replacement lost, the connection needed reconnecting from that moment. A provider can soften this with a few seconds of reuse grace, which is a decision to make with them, not a retry loop to build.
Rehearse a credential incident on the test client. The runbook's order is revoke, deploy, review:
Revoke: in Access Token Management, rotate
lab-tmp-printer-test's secret, as you would if the old one appeared in a public repository. Do not updatePRINTER_TEST_SECRETyet.Observe:
node printer.mjs test.json race avais refused with401 invalid_client; feed that to the policy withclassify refresh "$(printf '{"error":"invalid_client"}\n401')":alert_operators, never a retry.Deploy:
read -rs PRINTER_TEST_SECRETwith the new value, export it, and reconnect.Review: in Audit, filter by client
lab-tmp-printer-testand the last hour, and list what was issued with the old secret.
Note: the tenant has no overlap period for client secrets (G10), so the replacement must reach every running copy before the next request.
Rows the tenant cannot be made to produce on request.
Simulation. the tenant cannot be made slow or unavailable for a lab, so a local stand-in answers 503 and a closed port stands in for a connection failure. The policy code under test is your own, unchanged.
node -e 'require("http").createServer((q,s)=>{s.writeHead(503,{"Retry-After":"30","Content-Type":"application/json"});s.end("{\"error\":\"temporarily_unavailable\"}")}).listen(8799,"127.0.0.1")' &
classify refresh "$(curl -s -w '\n%{http_code}' -X POST http://127.0.0.1:8799/oauth/token)" # retry_after_delay
classify refresh "$(curl -s -w '\n%{http_code}' -X POST http://127.0.0.1:1/oauth/token)" # retry_with_backoff
kill %1
node --input-type=module -e "import { backoff } from './retry.mjs'; console.log([0,1,2,3,4].map(backoff))"
During an outage the connection stays connected but waiting, the calendar job reschedules itself, and nobody is told to reconnect, because their grants are fine.
Assemble the day's signals from real records. In Audit, filter Outcome: rejected for today. In Logs, open the protocol summaries. Count, with the first and latest request ID for each:
invalid_clientby client,invalid_grantby reason (code_replayed,refresh_replayed,refresh_revoked,pkce_failed), andrate_limited.
Note: there are no tenant metrics or alerts yet (G36), so these counts are assembled by hand. A jump in invalid_client right after step 3 is exactly what an alert should catch.
Check retention. Audit and Logs keep 30 days. Note the Logs coverage line for any summarized or suppressed requests, and write down how long you would keep operational records and security events, and why.
Break it
Steps 1 to 4 are the failures, each restored in place: the secret is deployed in step 3 and the stand-in stopped in step 4. Nothing else in the tenant is weakened.
Check your work
Press Check my progress before Cleanup, because Cleanup deletes the test client. The checks look for, in order: code_replayed from the retried exchange, refresh_replayed from the lost-response retry, the secret rotation, and the refused request with the undeployed secret.
Your retries.log has one tagged retry, and no line in it contains a code or token.
Cleanup
Make sure the stand-in on port 8799 is stopped.
Delete
retries.logand stored tokens:rm -f retries.log; rm -rf ~/lab-printer-oauth.Delete the client
lab-tmp-printer-test. Rununset PRINTER_TEST_SECRET REFRESH CODE VERIFIER.
Missing infrastructure
G36 Tenant metrics and alerts. Counts over time by operation, reason and client, with thresholds that notify an administrator, such as a jump in
invalid_clientafter a deployment or anyrefresh_replayed. Step 5 would then set an alert and see it fire after step 3.G10 Client secret rotation overlap. With a short period where both the old and new secret work, step 3 becomes a zero-downtime rotation rehearsal: deploy the new secret, confirm it is in use, then retire the old one.