writing

Start with the risk: how to prioritise security work

· 12 min read

A hiker walks carefully along a rocky mountain path beside a sheer cliff edge, with a valley far below

A security team never runs out of work. There are always gaps to close, controls to improve, tools that could be deployed, processes that could be better, and ideas the team would like to try. Some of that work comes from outside, like an audit finding or an engineer asking for an exception. But most of it comes from the team itself, because a good security team keeps finding things that could be better. The problem is that there are only so many people, and only so many hours.

So, how do you organise all that work? Almost everything on the list can be defended on its own, and that is exactly what makes it difficult. If you choose by what feels most urgent, by what is most interesting to work on, or by who is asking loudest, the team will be busy. But it may not be reducing the risks that matter most. And it will be difficult to explain to anyone, including the team, why some work was done and other work wasn't.

A good answer is to start with the risks that can really hurt the business, and work backwards to the controls that reduce them. It sounds obvious, and in a way it is. But I think it changes three things that matter a lot to a security team: how you plan the work, how you decide what not to do, and how you talk to the board.

Name the risks that matter

The first step is to write down a small number of high-level risks. By that I mean the events that would cause serious damage to the business if they happened. For most companies, the list will include ransomware, financial crime (fraud, payment abuse) and denial of service. Your list might have a few more, depending on what you sell and who you sell it to. A company that processes payments will look at fraud very differently from one that sells software licences, for example.

I think the most important thing here is to keep the list short. Why? Because the value of the list comes from forcing a choice. With thirty "top risks", almost any piece of work can be linked to one of them. A phishing simulation reduces one risk, a new logging pipeline reduces another, and hardening the build system reduces a third. Every piece of work now has a risk to justify it, so the list no longer helps you decide which one goes first. The discussion about which work matters most simply turns into a discussion about which of the thirty risks matters most, and you are back where you started.

With five or six risks, that is much harder to do. Some work will clearly reduce one of them, and some won't. That difference is what lets you organise the backlog.

A useful test is whether you can explain each risk to your CEO in one sentence. If you can't, it's probably too detailed for this level, and it may be better treated as one of the ways a bigger risk can happen. Credential stuffing, for example, is usually better seen as a route to fraud than as a top risk on its own.

Map the controls to each risk

Once you have the list, the next step is to write down, for each risk, the controls that reduce it. Take ransomware. The controls that matter most are usually offline backups that you restore regularly, MFA on every account that can reach production or the corporate network, endpoint detection and response, fast patching of anything facing the internet, and an incident response plan that the team has actually rehearsed. Notice that several of those depend on something being done regularly, not just being bought. A backup that has never been restored may not work when you need it, and you won't find out until then.

With this map in place for every risk, prioritisation becomes much easier. When a new piece of work comes up, whether it's the team's own idea or a request from someone else, you ask one question: which of our top risks does this reduce, and by how much? If the answer is "none of them", it goes to the back of the queue. And, importantly, you can explain why in a sentence. The conversation is no longer about whether the work is a good idea, because it may well be. It is about whether it is more important than the work that reduces the risks we have agreed matter most.

The map also shows you the gaps. A big risk with only one or two weak controls against it is where the next pound (and the next hire) should go. It works the other way round too. A control that doesn't map to any risk deserves a closer look, because someone is paying for it, in licences, in maintenance or in the time of the people who run it.

Inherent risk and residual risk

There are two terms that help here: inherent risk and residual risk. Inherent risk is the risk before any controls are applied. In other words, how likely an event is, and how bad it would be, if nothing stood in the way. Residual risk is what is left once the controls are in place and working.

Take ransomware again. Without backups, MFA or EDR, the inherent risk may be that an attacker encrypts most of the company and there is no way back. With offline backups that we restore regularly, the same attack still hurts, but we can recover. That difference is what the backups are worth.

The same applies to every control. That is why, when comparing two pieces of work, the question to ask is how much each one lowers the residual risk on one of the top risks. Residual risk is also what matters for decisions. It's what we compare against the level of risk the business is willing to carry, and it's what we formally accept when we decide to stop there.

And notice the words "and working" in that definition. They are easy to read past, but they are important. Both types of risk move over time. Inherent risk goes up when the threat grows or the business changes: more payment volume, for example, means more for fraudsters to go after. Residual risk goes up when a control quietly stops working, which brings us to the next point.

Check your controls still work

Controls age, and threats move faster than most control reviews. A control that was effective two years ago may be doing very little today, and nobody notices because nothing is visibly broken. There is no alert for a control that has stopped being relevant.

Captchas are a good example. For years, they were a decent way to stop bots from creating fake accounts or trying stolen credentials at scale. Today, AI-driven browser automation can solve many captchas cheaply and quickly, so attackers get through while real customers still pay the cost in friction. The captcha is still there, and the dashboard still shows it as "in place". But is it still reducing the risk we bought it for? If it isn't, the residual risk has gone up, even though nothing on the control list has changed. And it's slightly worse than that, because the list now tells us we are safer than we are.

That is why it is worth evaluating every control on a regular schedule, against the threats we see today rather than the ones we saw when we deployed it. A few simple questions help:

The last question is the most useful one. If nobody can answer it, we probably don't understand what the control is doing for us. And if we don't understand that, we can't really say how much it lowers the residual risk either.

Write down the risks you accept

You won't fix everything, and that's fine, as long as the decision is deliberate and written down. What we accept is always the residual risk. When a risk is accepted, the record should say what the risk is in plain words, why it is being accepted (cost, timing, low likelihood), who agreed to it, and when it will be looked at again.

Who agrees to it matters. The business owns the risk, so the acceptance has to be agreed with the business and with the people who own the technical controls that mitigate it. Security's role is to make sure everyone understands what they're signing up to. That means describing the risk in terms the business recognises, not just naming a missing control.

A written acceptance also protects everyone involved. If the risk turns into an incident six months later, nobody has to argue about who knew what, and the conversation can move straight to fixing it. The review date matters as much as the signature. A risk that was acceptable last year may look very different once the threats change, and the captcha example above is exactly that.

Give the board what the board needs

Board time is short, and a long list of everything the security team has done is not a good use of it. It isn't relevant to what the board has to decide either. Instead, it helps to focus on two things:

  1. Material risks. These are the risks that could cause serious financial loss, regulatory action or permanent closure. Ransomware sits here for most companies, and most companies are not adequately prepared for it. The board needs to know where we stand against each material risk, whether that position is getting better or worse, and what we need from them.
  2. Emerging risks. These are the risks that weren't on the register last year. Geopolitical risk is a good one to bring up now. It can feel abstract in a board meeting, so it helps to use something that has already happened. The drone strikes on AWS's data centres in the UAE in March 2026 are a good way to make it concrete.

What March 2026 taught us about losing a region

On 1 March 2026, one of the three Availability Zones in AWS's UAE region (ME-CENTRAL-1) lost power. At first, AWS described it as a localised power issue. A few hours later, it gave more detail: objects had struck the data centre and started a fire, and the fire service had cut power to the building, generators included, while they put it out. Early the next day, a second Availability Zone went down. At that point, customers couldn't launch new instances anywhere in the region, and core services like S3 and DynamoDB were failing a significant number of requests. AWS then confirmed what had happened. Two of its facilities in the UAE had been directly hit by drone strikes, and a facility in its Bahrain region had been damaged by a strike close by.

I think these details are worth staying with for a moment, because they explain why this was so different from a normal outage. A region is designed to survive the loss of one Availability Zone. It is not designed to survive the loss of two at the same time. The third zone in the UAE kept running, but even there some services were affected, because they depended on the zones that were down. AWS later said about Bahrain that the damage had gone beyond what its regional and multi-AZ services were built to cope with, and the same was clearly true in the UAE.

And then there is the recovery. With a software bug, you wait for a fix. Here, within a day, AWS was telling customers to enact their disaster recovery plans and restore from backups into other regions. By the end of April, it said the UAE region couldn't reliably support customer applications, and that restoring normal operations would take several months. In September, it confirmed that it couldn't restore the resources and data held only in the first zone that was hit. Most customers had moved to other regions by then. But anything that existed only in that zone is, for practical purposes, lost. The full sequence is in AWS's own updates on its Health Dashboard.

So, what does this mean for a board? I think it is the clearest example we have of geopolitical risk to infrastructure. A region we depend on becomes unavailable, and so does everything we didn't realise depended on it. And the usual assumption, that the provider will fix it and we just need to wait, doesn't hold. The questions change. Do we have backups outside the region? Have we ever restored from them? And where are we allowed to move to?

That last question isn't only technical. When AWS suggested alternative regions, it pointed customers to the US, Europe or Asia Pacific, depending on their latency and data residency requirements. If regulation says our customers' data has to stay in a particular country, the options may be much narrower than our architecture diagram suggests.

Concentration in a single cloud region is a risk like any other. It goes on the register with an owner, a set of controls and a review date, and it gets the same questions as everything else: which controls reduce it, are they working, and how much of it are we willing to accept?

A word on heat maps

Heat maps are useful. They give the board a picture of many risks on a single slide, and I understand why people like them. But, as Eric Staffin observed in a LinkedIn comment, a risk can move across a heat map while the real exposure stays exactly where it was. Well, in fact, sometimes the risk doesn't even move: the heat map is redrawn with more squares, and the movement appears on its own. Either way, the board gains nothing from it.

What the board needs to know is whether leadership decisions are driven by risk, whether those decisions are reducing what the business stands to lose, and whether the worst outcomes stay within the level of risk the organisation has agreed to carry.

So, reporting progress on a material risk needs to be backed up with evidence that the exposure has changed:

The board can then judge the progress for itself. The colour on the slide is still there, but it becomes a summary of changes they can see, rather than the only thing they are asked to trust.

Doing it consistently

None of this is complicated on paper: a short list of risks, a map of controls, regular testing, written acceptances, and a board report that shows what has actually changed. The difficult part is doing it consistently, quarter after quarter, while the list of things the team would like to do keeps growing.

But that is exactly when this approach is most useful. There will always be more work than people. The risk list is what lets you decide what gets done, and explain to the team and to the rest of the business why some work has to wait.

All writing