OPENID CONNECT · LAB
Reproduce and diagnose seven sign-in failures
Produce a failure at each stage of a sign-in with real wrong inputs, decide the stage from where the browser ended up, and find the matching tenant record, including a server error found by its reference_id.
ReadyUses your lab tenant
The lesson
Builds on: Implementing a relying party.
New to the labs? Start with the lab toolkit and the shared cast and names every lab uses.
Your progress
Press Start before you begin. Only events your tenant records after that count, in the order below. Checking reads your tenant's Audit, so you need Audit read access in it.
Sign in to start this lab and check your progress. Log in or create an account.
Rotate lab-collage's client secret
Recorded as
tenant.oauth.credentials.rotatesucceeded.Present the old secret at the token endpoint
Recorded as
oauth.tokenrejected (invalid_client) forlab-collage.Redeem one code twice
Recorded as
oauth.tokenrejected (code_replayed) forlab-collage.Send repeated silent requests with no session
Recorded as
oauth.authorizerejected (login_required) forlab-collage.Trigger a server error with a reference_id
Recorded as
oauth.authorizefailed (id_token_policy_unavailable) forlab-collage.
Setup
Each failure is produced with real requests and real wrong inputs, and each leaves evidence in the tenant's Audit or Logs, or in your own handler's log.
Use the handler,
beginandpending.jsonfrom Build a callback handler that fails closed.Give your relying party a log line for every failure, with a fixed message, the stage, the check, the issuer, a
kidwhen there is one, and a correlation ID, never a code or token:
rp_log() { jq -cn --arg stage "$1" --arg check "$2" --arg iss "$ISSUER" --arg kid "${3:-}" \
'{time: (now | todate), message: "Sign-in failed during \($stage)", stage: $stage, check: $check, issuer: $iss, kid: (if $kid == "" then null else $kid end), correlation_id: ("corr-" + (now | tostring))}' | tee -a rp-log.jsonl; }
source ~/btl-oidc.sh, runbtl-lab callbackbefore each request, and press Start.
Walkthrough
An error page at the provider. Send a redirect URI with a trailing slash:
REDIRECT_URI="http://127.0.0.1:8765/callback/" signin
The tenant shows its own error page and the listener prints nothing. Logs counts oauth.authorize rejected invalid_redirect_uri. Compare what you sent with the registered value character by character, then log rp_log request invalid_redirect_uri.
Why it matters: a provider never redirects, even with an error, to an address it does not recognize, so the relying party sees only an attempt that never came back.
invalid_clientafter a rotation. In OAuth > Clients > lab-collage, rotate the secret. Sign in, then exchange the code with the secret you still hold:
redeem '<code>'
{"error":"invalid_client"} with status 401. Store the new secret with read -rs CLIENT_SECRET and log a version label for it, never the secret: rp_log token_exchange invalid_client.
Why it matters: in this tenant the old secret stops at once (G10), so every server must switch together. If only some requests fail, look for the server that missed the rotation.
invalid_grant. Redeem one code twice, then ask about the first access token:
redeem '<code>'; FIRST=$TOKEN
redeem '<same code>'
curl -s -u "$CLIENT_ID:$CLIENT_SECRET" "$ISSUER/oauth/introspect" --data-urlencode "token=$FIRST" | jq .active
The second answer is invalid_grant, and the first access token is now false.
Why it matters: a second presentation of a code may be an attack or a retry bug. The provider revokes the whole family either way.
An unknown key. If you kept
cached-verify.mjsfrom Rotate a signing key against a cached key set, generate a new ID token key, point the manager at it, sign in, and run the check withNO_REFETCH=1and then without. Log the token'skidand the identifiers in your cache:rp_log id_token_validation unknown_kid "$(part "$ID_TOKEN" 1 | jq -r .kid)".
Why it matters: one refetch fixes rotation day. A refetch that still lacks the key points at the token's issuer.
A nonce mismatch from your own bookkeeping. With the plain helpers, which keep one global
NONCE, runsignintwice, open both URLs in two tabs, and complete the first:
btl-lab verify "$ID_TOKEN" --issuer "$ISSUER" --audience "$CLIENT_ID" --type id --nonce "$NONCE"
It fails on the nonce, because NONCE now belongs to the second attempt. Repeat with begin and callback: both tabs succeed.
Why it matters: most nonce failures are client bookkeeping, not attacks.
A
prompt=noneloop.curlcarries no tenant cookie, so it plays a client with no session that answers eachlogin_requiredwith another silent request:
for i in 1 2 3 4 5; do curl -s -o /dev/null -w '%{redirect_url}\n' "$(signin prompt=none)"; done
Five callbacks with error=login_required, and Audit shows five rejections in seconds. Write the limit into your design: one silent check per visit, at most three sign-in attempts per browser in five minutes, then a page that explains what happened.
Why it matters: login_required is an answer. A loop needs a limit as well as a fix.
A server error found by reference. In OAuth > ID token managers, create
lab-tmp-broken-id(not the default), assign it tolab-collage, then disable it. Runsignin:
signin
The listener prints error=server_error, an error_description, and a reference_id. In Logs, search for that reference_id: oauth.authorize failed id_token_policy_unavailable.
Why it matters: the tenant tells the client nothing internal, yet gives it a value that leads support straight to the cause.
Break it
The wrong fix for clock skew. Make your own time check take a tolerance, and set the clock five minutes behind:
times_ok() { local now=${NOW:-$(date +%s)} tol=${TOLERANCE:-30}; part "$1" | jq -e --argjson now "$now" --argjson tol "$tol" '.exp > ($now - $tol) and .iat <= ($now + $tol)' > /dev/null && echo ok || echo "FAIL iat in the future"; }
NOW=$(( $(date +%s) - 300 )) times_ok "$ID_TOKEN"
NOW=$(( $(date +%s) - 300 )) TOLERANCE=600 times_ok "$ID_TOKEN"
The first fails on iat; the second passes because the tolerance swallowed the drift, and every other time check got weaker with it.
Restore: keep the tolerance at 30 seconds. The real fix is synchronizing the server's clock.
Check your work
Press Check my progress. The checks look for the secret rotation, the old secret refused, a code redeemed twice, the silent loop's login_required answers, and the server error recorded with its reason.
In Logs, find invalid_redirect_uri in the protocol summary and the failed record by its reference_id. Your rp-log.jsonl has one line per failure, each with a stage and check and no secrets.
Cleanup
Assign
lab-collageback to its previous ID token manager and deletelab-tmp-broken-id.Keep the rotated secret in
CLIENT_SECRETonly.Delete
rp-log.jsonlandpending.jsonwhen you are done with the track. To start the OpenID Connect labs again from a clean state, reset the lab tenant with the Lab Photos preset.