Layered defenses and secure defaults
No single control
Cedar Inc. did not lose control of Riley's account for lack of security. Riley, who works in Cedar's finance team, signed in with a password and an authenticator-app code. Even so, one morning an attacker relayed Riley's sign-in through a look-alike page, kept the session, added their own authenticator app, approved a mail app called "Invoice Sync", set up a forwarding rule, and used the app to read a payments vendor's API key from Riley's mail. An alert caught it the next morning.
Quinn, who led the investigation for Cedar's security team, asks two questions about every step: what would have stopped it here, and if that had failed, what would have stopped it at the next step? That is the idea behind defense in depth: several layers of protection, arranged so that an attacker who gets past one still meets another. A layer can stop a step outright, or limit it by making the step slower, smaller, or visible to someone who can respond. Some of Cedar's layers did their job. At other steps, nothing stood in the way.
| Step | What happened | What was in place |
|---|---|---|
| Day 0 | Common passwords were tried against hundreds of Cedar accounts. | Sign-in limits. Held: nothing succeeded. |
| Day 1, 09:14 | Riley signed in through the relay. | A password and an authenticator-app code. Both passed through. |
| 09:20 | The attacker added their own authenticator app. | A recent sign-in check. Passed. |
| 09:31 | The attacker approved "Invoice Sync" to read and send mail. | Nothing. Any employee could approve any app. |
| 09:40 | A rule began forwarding messages containing "invoice" outside Cedar. | Nothing. Any mailbox could forward mail anywhere. |
| 10:05 | A caller asked the help desk to reset an administrator's MFA. | A callback procedure. Held: refused and recorded. |
| 11:30 | Invoice Sync read a vendor email containing an API key. | Nothing. The key had been sent by email. |
| Day 2, 08:50 | An alert combined three weak signals. | Detection. Held, almost a day later. |
Before reading on, try matching each step that got through to the defense that would have stopped it.
Where would it have stopped?
Simulation. These are the steps of the made-up Cedar incident that succeeded. For each one, pick the defense that would have stopped it. Answering sends nothing anywhere.
Layers placed early do the most good. Had Riley signed in with a passkey, the browser would not have offered it at the look-alike address, and none of the later steps would have happened, as Phishing and relayed sign-ins explains. Early layers have gaps, though. If Cedar's policy still let Riley fall back to a code when a passkey "did not work", the relay page would simply ask for the code, and the later layers would be what made that gap survivable.
Layers also count as separate only if they fail for different reasons. Cedar's recent sign-in check relied on the very session the attacker had just stolen, which is why it passed, as Limiting what stolen access can do explains. A passkey confirmation at the same moment relies on something a relay cannot carry, which makes it a genuinely separate layer.
Defaults people never change
Two of Cedar's steps succeeded without the attacker defeating anything. Cedar's configuration let any employee approve any app, whatever access it asked for, and let any mailbox forward mail anywhere. Nobody had weighed those settings. They were what the systems came with.
Most settings stay at their defaults, because administrators are busy and a setting that has never caused trouble looks fine. A secure default is a starting setting that is safe without anyone acting on it: apps that ask for broad access wait for approval, forwarding outside the organization is off, administrator roles expire, and new accounts enroll a strong method before they can do much. Someone who genuinely needs a looser setting can still have it, but they have to ask, and the request leaves a record.
Exceptions need the same care. A rule that exempts a team from passkeys "until the new laptops arrive" quietly becomes permanent once the laptops arrive and nobody removes it. An exception with an owner, a reason and an end date stays visible, and lapses unless someone renews it. Without those, it becomes one of the paths that never ask for the evidence the policy requires.
People who build identity software make the same choice for everyone who uses it. If a token library checks the signature, issuer, audience and expiry unless told otherwise, nearly every application that uses it will check them. If those checks have to be switched on, many applications never switch them on, and validation that lets forgeries through becomes common.
Failing safely
Every check can fail to run. A policy service times out, signing keys cannot be fetched, a risk score never arrives. The system still has to answer the request in front of it, and its answer when it cannot decide is as much a security decision as any rule it enforces.
A system that fails closed refuses a request it cannot evaluate. One that fails open allows it. Failing open keeps people working during an outage, but it hands an attacker a new move: make the check fail, or wait until it does, and the request goes through. Some failures can be caused on purpose, by overloading a dependency or by sending input that makes the check throw an error. Code often fails open by accident, in an error handler written for convenience:
Fails open
try:
allowed = policy.check(user, action, resource)
except PolicyUnavailable:
allowed = True # an outage grants access
Fails closed
try:
allowed = policy.check(user, action, resource)
except PolicyUnavailable:
record_event("policy_unavailable", action)
allowed = False # no decision means no access
Failing closed has a cost of its own, so choose deliberately for each check rather than leaving it to whatever the code happens to do. An application might keep using cached signing keys for a short time when it cannot fetch fresh ones, and raise an alert, because those keys rarely change. A sensitive action, such as adding a sign-in method, waits until the check works again. For the day the identity provider itself is unavailable, organizations prepare emergency access accounts in advance instead of switching checks off.
Processes can fail safely too. At 10:05, Cedar's help desk refused an MFA reset it could not confirm through its callback procedure, the call described in Attacks on recovery and the help desk.
Checking every request
At 09:20, Cedar's identity provider received a request to add an authenticator to Riley's account. The request carried a valid session cookie, so on that point it looked like Riley. Other facts disagreed. Riley normally signs in from Cedar's office network, 198.51.100.0/24, and this request came from 203.0.113.57, a hosting network Riley had never used, minutes after a sign-in that had also come from a hosting provider. A system that asks only whether the session is valid accepts the request. A system that looks at the whole request has reasons to ask for more.
Older designs drew a line around a network or a sign-in and trusted whatever was inside it. The alternative is continuous verification: each request is evaluated with current information, such as who is asking, how and when they signed in, which device and network the request comes from, how sensitive the action is, and what the person's roles allow right now. Most of this happens without the person noticing. A routine request with familiar signals goes straight through, while adding a sign-in method from an unfamiliar hosting network, in a session established with codes a relay can pass along, calls for a passkey confirmation or a refusal.
Zero trust is a widely used name for this approach. No request is trusted because of where it comes from or because an earlier check passed, and access is decided per request with current information about the person, the device and the resource. NIST SP 800-207 describes the approach for an organization's networks and systems. It is not a product, and it does not mean distrusting the people who work there. Each decision is made with what is known at that moment, instead of being inherited from a check made earlier.
That only works if the information is current. A role removed at 10:00 has to stop granting access at 10:00, not when a cached token expires an hour later. This is why permissions are checked on the server for every request, and why identity providers and applications share security events such as a revoked session or a disabled account.
Security people can live with
Whether people can live with a control is part of whether it works. A control that people find painful collects exceptions, workarounds and habits that defeat it. Send an approval prompt for every sign-in and people learn to approve without reading, the habit described in Getting around MFA. Force a new password every month and passwords become predictable variations of each other.
A passkey is quicker to use than a password and a code, and it is the method that would have stopped the relay at 09:14. Asking for stronger evidence only where it matters, such as adding a sign-in method or approving an app with mail access, keeps routine work smooth and puts the friction where an attacker has to pass.
People are part of the defense too. A simple way for Riley to report an odd email or an unexpected prompt, and a response that thanks people for reporting, turns everyone who uses the system into another layer of detection.
Layers and defaults that are right today drift. Exceptions pile up, new apps arrive, and someone who gained administrator access during an emergency keeps it. Reviewing identity security posture covers checking that the layers are still where they were put.