IDENTITY SECURITY · LAB
Write detection rules and test them against your own Audit
Produce a small, harmless pattern of failed sign-ins and a quick method change, then write detection rules in plain words, find their matches by hand in Audit and Logs, and decide each rule's response.
Partly readyUses your lab tenant
The lesson
Builds on: Recording identity activity.
New to the labs? Start with the lab toolkit and the shared cast and names every lab uses.
Partly ready. Most of this lab runs today. Steps that wait on platform features are marked, and Missing infrastructure says what they need.
- G36 Detection, risk signals, tenant metrics and alerts (baselines, anomaly rules, thresholds that notify an admin)
- G37 Security context in Audit (network, device, session, method; subject on failed sign-ins) and user "this wasn't me" reporting
Your progress
Press Start before you begin. Only events your tenant records after that count, in the order below. Checking reads your tenant's Audit, so you need Audit read access in it.
Sign in to start this lab and check your progress. Log in or create an account.
Mistype passwords for several lab people
Recorded as
account.sign_inrejected (invalid_credentials).Cora signs in
Recorded as
account.sign_insucceeded about[email protected].Cora adds a sign-in method minutes later
Recorded as
account.securitysucceeded (method_enrolled) about[email protected].Read the protocol Audit to test the rules
Recorded as
tenant.protocol.audit.listsucceeded.
Setup
Press Start on this page.
Open a notes file with one section per rule: the rule in plain words, the matching records with their request IDs, what makes it noisy, and the response it should trigger.
Walkthrough
Produce a small pattern. At
$ISSUER/login, type one wrong password each for Ava, Ben and Cora, a few seconds apart. Then sign in as Cora correctly and, within a few minutes, add an authenticator app at$ISSUER/account/security(skip any enrollment prompt and use the Security page).
Why it matters: one failure means nothing, and one new method means nothing. The lesson's point is that detection looks for patterns across accounts, time and kinds of events. You now have both a spread of failures and a sign-in followed quickly by a new method.
Rule 1, failures spread across accounts: "more than N rejected password sign-ins across different accounts within one hour from the same network." In Audit, source Protocol activity, filter
rejectedand count theaccount.sign_inevents withinvalid_credentialsin the last hour. Then compare with the day's count in Logs, source Protocol summaries.
Why it matters: this is the spraying shape: few attempts per account, many accounts. Note what you cannot see: the records hold no subject and no network for these failures, so you can count them but not group them by account or source.
Rule 2, a new method soon after a sign-in: "a sign-in method added within 10 minutes of a sign-in, by the same account." Find Cora's
account.sign_inandaccount.securitywithmethod_enrolled, and write both times and request IDs.
Why it matters: this is Cedar's combined signal with two of its three parts. The third part, a network the account has never used, is missing from your records, which is why the rule here is weaker than Cedar's.
Rule 3, refused attempts to gain more: "any refused role edit, role assignment or methods reset." In Audit, source User directory, filter
rejectedand list the refusals from the escalation and recovery labs with their actors.
Why it matters: refused requests show intent. An ordinary account trying to give itself more is rarely an accident, and the actor field names who tried.
Rule 4, help desk resets of privileged accounts: "a methods reset whose subject holds a management role." List the
tenant.users.methods.resetevents and check each subject against your privileged-access list from the privileged access lab.
Why it matters: the lesson lists requests to reset MFA for administrators as a signal. In your tenant the reset itself is recorded, so the rule can compare it with who holds power.
For each rule, write what makes it noisy and the baseline that would quiet it. For example, Rule 2 fires for every new user who sets up a method after their first sign-in.
Why it matters: unusual is not the same as malicious. A rule without a baseline trains its readers to dismiss it, and that is how a real alert goes unread.
For each rule, choose a response that matches confidence and stakes: record only, ask for stronger sign-in, notify the owner on a channel the attacker is unlikely to control, hold the change, or block and end sessions. Note which of these you can do by hand in your tenant today: lock the user, reset methods, revoke tokens.
Why it matters: the lesson's responses run from gentle to drastic. Blocking every unusual sign-in locks out legitimate people; asking for a passkey costs seconds.
Write Rule 2's match as an alert in the lesson's format: a plain summary, the account, evidence lines with times and request IDs, the baseline (or "unknown"), and the next step.
Why it matters: an alert is the reader's starting point. One that carries its evidence and the next action lets someone act in minutes instead of rebuilding the context.
Break it
Apply Rule 2 to Mia, who set up an authenticator again after the recovery lab. It matches, although nothing is wrong. Context, a baseline, and confirming with the owner on another channel decide; the rule alone cannot.
Check your work
Check my progress confirms the failed sign-ins, Cora's sign-in, her new method a few minutes later, and your Audit review.
Your notes contain four rules, each with matching request IDs, a noise source and a chosen response, plus one alert written in the lesson's format.
Cleanup
Leave Cora's authenticator app in place, or remove it at
$ISSUER/account/security.
Missing infrastructure
G36: the tenant runs no rules, baselines, risk scores or alerts. You evaluated each rule by hand. Once it exists, the lab will enable Rule 2 as a tenant detection, repeat step 1, and measure how long the alert takes to arrive.
G37: failed sign-ins carry no subject, and no record carries a network, device or session ID, so Rule 1 cannot group by account or source and Rule 2 cannot tell a new network from a familiar one. Once it exists, the lab will repeat Rule 1 and group the failures by source.