Testing identity defenses
Testing what should fail
Cedar Inc. had tests for the screen where employees add a sign-in method. They checked that someone who had signed in recently could add an authenticator app, that the new app produced working codes, and that it appeared in the person's list of methods. Every test passed. Then an attacker relayed the sign-in of Riley, who works in Cedar's finance team, through a look-alike page, kept the session, and used that screen to add an authenticator of their own. The recent sign-in check passed, as it always will for a session stolen minutes earlier. Nothing in Cedar's tests had asked what the screen should refuse.
Most tests confirm that the right thing works: the correct password signs in, and the authorized person can export the report. Attacks use the other paths, so identity systems also need negative tests, which send a request that should be refused and confirm that it is, as the deny tests in Testing, auditing, and changing policy do for permissions. For the screen where a sign-in method is added, they might look like this:
| Case | Expected result |
|---|---|
| No session, or an expired one | Refused, and the person is asked to sign in. |
| A session older than the recent sign-in window | Refused until the person confirms with a passkey or security key. |
| A fresh session established with codes, where the policy requires a passkey for this change | Refused until a passkey confirms it. |
| A valid session for one person, with another person's account ID in the request | Refused, and recorded as a rejected attempt. |
| The same confirmation sent twice | The second one refused. |
| A request sent directly to the API, without the page | The same checks as through the page. |
| Many attempts in a short time | Slowed or refused by limits. |
A refusal is worth checking for what it says and what it leaves behind. The response should reveal no more than it needs to, such as whether an account exists, and the refusal should appear in the records with an actor, a subject and a reason, as Recording identity activity describes.
Negative tests earn their keep when they run automatically with every change. A permission check that disappears during a rewrite, while the button stays hidden from ordinary users, passes every test that only exercises the button, which is why checks run on every request and tests should send requests directly. The other ways into an account need the same attention. Recovery, older protocols and the help desk's reset procedure all lead to the same account, and every path to it needs tests of its own.
Permission comes first
Testing defenses means doing what an attacker would do: guessing passwords, sending altered requests, calling the help desk with a made-up story. Done without permission, that is an attack, whatever the intention. It can break the law, lock real people out of their accounts, set off a real incident response, and expose other people's data. So every test, from a single negative test to a penetration test that looks for weaknesses across a whole system, starts with written authorization from someone who has the authority to grant it for that system, and with an agreed scope that says exactly what may be tested, how, and when.
Test only systems you are authorized to test. Having an account on a service does not authorize testing the service. A colleague's spoken go-ahead does not authorize testing a system that colleague does not own. An organization's permission also covers only what it controls. If Cedar's sign-in runs on another company's identity service, or one of its applications is hosted by a vendor, those parts need the provider's own permission or stay out of scope, and many providers publish the rules under which they allow testing.
A scope written before testing starts answers practical questions: who approved it, which systems and accounts are included, what is excluded, when testing happens, when to stop, and who to call. Here is a fictional summary for a test at Cedar:
Authorized by Cedar Inc. head of security, owner of login.cedar.example
Testers Cedar security team
When Five working days, 08:00 to 18:00 Cedar time
In scope Sign-in and sign-in method changes at login.cedar.example,
using test accounts created for this test only
Help desk reset procedure: two calls, agreed in advance
Out of scope Real employee and customer accounts
The payments vendor and every other third-party service
Overloading services or locking out real people
Stop and report Signs of real attacker activity, or real personal data
Contact Security on-call, who can confirm the test is authorized
Results Shared only with the security team and system owners
Tests that involve people, such as calls to the help desk, need their own explicit approval and an agreement about how the staff involved are treated afterward. The aim is to test the procedure, not to catch someone out. A help desk agent who is fooled during a test has shown where the procedure needs support, and blaming them teaches everyone else to hide mistakes. Throughout, use accounts created for the test rather than real people's accounts, and stop as soon as anything real turns up.
Testing detection
At Cedar, the attacker added an authenticator at 09:20 on the first day, and the alert that caught it fired at 08:50 the next morning. Controls that refuse requests are only half of a defense. The other half is noticing what gets through, and that needs testing as much as any refusal does.
A detection test starts from an action whose result you can predict. Within the agreed scope, sign in to a test account from a network it has never used, then add a sign-in method. Then follow what should happen next. Did the records capture the sign-in and the change, with the right actor, subject and network? Did the alert fire, and how long did it take? Did it reach someone who could act, with enough information to act on, as Detecting identity attacks describes? Did the account's owner get a notice? Each answer is a measurement, and comparing measurements over time shows whether detection is getting faster or quietly breaking.
Detection usually breaks without a sound. A log source changes its format, a field is renamed, an integration stops sending events, and the rule that depended on them never fires again. Nobody notices an alert that does not arrive. Running the same detection tests on a schedule catches those failures before an attacker benefits from them. The same tests can confirm that the records hold nothing they should not: no passwords, codes or session cookies among the fields, as Recording identity activity explains.
Decide in advance whether the people who watch alerts know a test is happening. Telling them avoids a real response to a test. Not telling them tests their response as well, but then someone with authority must be able to confirm quickly that the activity is authorized, which is why the scope names that contact. An unannounced exercise in which testers act out a realistic attack from start to finish is usually called a red team exercise. When testers and the people who watch alerts work side by side instead, checking after each step what the records and alerts showed, the practice is often called purple teaming, and it is one of the quickest ways to tune detection.
Rehearsing an incident
Some parts of a response cannot be tested by sending requests, because they depend on people knowing what to do and being able to do it. A tabletop exercise rehearses those parts. The people who would respond to a real incident sit together and talk through a fictional scenario, step by step, without touching any system. A facilitator describes what happens and asks what each person would do next.
At Cedar, the scenario might begin: "It is 08:50. An alert says a new authenticator app was added to a finance employee's account from a network the account had never used, minutes after a sign-in from a hosting provider." Then the questions start. Who can end the employee's sessions in every application, and how long does that take? Who can withdraw an app's access to a mailbox? How do you reach the employee through a channel the attacker cannot read? Partway through, the facilitator adds a complication: the password was reset, but the attacker is back an hour later. The group has to work out that the attacker's authenticator was never removed, the gap Containing and recovering warns about.
Invite everyone who would be involved for real: the security team, the help desk, the owners of the affected applications, and whoever would speak to vendors or to employees. The most useful results are the gaps the conversation exposes: a step only one person knows how to do, a permission nobody on call holds, a vendor contact that lives only in someone's old email. Record each one with an owner and a date, the same way posture findings are fixed.
Testing recovery
Recovery is the part of a response most likely to be done for the first time during a real incident, which is the worst time to learn it. Almost every step can be practiced on test accounts and test systems first.
Practice giving an account back: verify the owner through a channel the attacker never controlled, remove every method and grant the owner did not add, and hand over a working account. Practice ending access too. Revoke a test account's sessions and refresh tokens, then measure how long each application keeps accepting what it already holds. That delay, described in The gap before a token expires, is far easier to measure in a quiet week than in the middle of an incident.
Rotate secrets and keys on a schedule, so the procedure is familiar when a leak forces it. A planned rotation reveals what depends on a secret in calm conditions, instead of through an outage, as Rotating credentials and keys after exposure explains. Cedar had to replace its payments vendor's API key during its incident. A team that has rotated that key before already knows where it is used and whom to call at the vendor.
Test emergency access accounts as well. On a fixed date, sign in with one, confirm that it still works and that its use raised an alert with the right people, then replace its credentials and store the new ones safely, as after any use. One test of that kind checks two things at once: the account that is meant to work on the worst day, and the alert that is meant to notice when someone uses it.
The Cedar incident makes a ready-made test plan. Within an agreed scope, Cedar's security team replays each step the attacker took against a test account: a sign-in from a hosting network, a new authenticator minutes later, an app asking to read and send mail, a rule forwarding mail outside Cedar. Each step should now be refused, held for a passkey or an administrator's review, or reported by an alert within minutes rather than the next morning, and the team then practices the cleanup on the same account. Whatever still gets through answers the last of the four questions from Threat modeling an identity system, did we do a good enough job, and goes back into the model as the next thing to fix.