Alan Simpson at Rapid7 describes how organisations can correctly validate that their security tools are working

Ask a board to name the organisation’s biggest cyber-risks, and you’ll usually get a sensible answer of ransomware, nation-state activity, or supply chain compromise. Ask the slightly different question of how you know your organisation could stop them, and things become less certain.
That is the gap purple teaming and security validation close. They expose the assumptions everyone has stopped questioning, such as a telemetry feed believed to be reaching the SIEM but was not, or a detection that stopped firing without anyone noticing.
No control stays reliable permanently in an environment that changes this fast, but boards rarely ask for that proof directly. They ask whether the organisation is secure, get a reassuring answer, and move on. The harder, more useful question is whether that reassurance was ever actually tested.
A purple team exercise only matters once findings travel beyond the people who ran it. The value comes from understanding which controls failed, which assumptions were wrong, and which gaps create meaningful business exposure.
I have seen a more extreme version of this first-hand. In a previous role, a red-team exercise was commissioned without the defensive team’s knowledge. During the engagement, close to 700 alerts were generated, two compromised machines were contained, and the phishing element failed until somebody was specifically asked to click.
Allowing the exercise to continue after a control succeeds can be legitimate if you need to test later stages of the attack path. The problem came when results were presented upwards. The story became that the testers had effectively walked into the environment without generating meaningful alerts, and that was not what had happened.
If you allow an exercise to progress beyond controls that worked, the board-level narrative still has to acknowledge that those controls worked. Otherwise, useful evidence becomes a “gotcha” exercise. Red teaming and purple teaming are there to expose weaknesses and test assumptions, not prove that one team is cleverer than another.
Purple teaming should make that easier. Put offensive and defensive teams together so they see the same activity and understand what each side can and cannot see. If something passes through without detection, finding that gap is the exercise working. When something is detected, contained, or blocked, that should be part of the evidence too.
A board doesn’t need packet captures or a walkthrough of every technique. It needs to understand which exposures materially affect the business, how confident the organisation is that its controls would prevent them, and where the next pound of investment will reduce risk most.
The output should not be another report presented once and forgotten until the next audit cycle. It should be a prioritised view of tested exposures, with enough context to show which ones matter.
Vulnerability severity on its own tells you very little about business risk. Imagine two findings. The first carries a severe headline score, but an attacker would need to bypass several controls before reaching it. The second looks relatively minor but sits directly on an attack path an intruder could discover within minutes. A generic severity model may prioritise the wrong one.
That might produce a better-looking vulnerability graph, but it does not make the organisation safer. Security teams can become very good at measuring activity, but vulnerabilities closed and alerts investigated are not the same thing as evidence that risk has been reduced.
The board should be pushing on three things: which exposures are genuinely exploitable, which controls held up and which didn’t, and how long ago anyone checked. An annual pen test is honest about one moment in time. It’s a lot less honest a few months later.
Someone spins up a new cloud service. Someone else picks up admin rights for a project and just never hands them back. None of that waits for the next scheduled test, so by the time it comes round, you’re not really testing the organisation you think you are.
That’s why cadence matters as much as prioritisation. Fix what’s already been proven to matter while there’s still time, before someone else finds it first.
A green dashboard feels like proof. It isn’t. An alert can fire exactly the way it was built to and still achieve nothing if nobody’s actually watching the queue it lands in. Having a control in place and knowing it works are two different things, and most security programmes can only really prove the first one.
Controls also decay. Permissions accumulate, configurations drift, and organisations still place confidence in the fact that something was checked once.
Validation also has to go beyond replaying a known attack against an endpoint. Breach and attack simulation has value, but there is a limit to what you learn from simulations that begin from an assumed foothold. You may prove that an alert fires once somebody is on the machine without proving whether they could have reached it in the first place.
This is where human experience matters: a skilled tester does not just ask whether a technique triggered an alert, but what comes next. That kind of adversarial thinking is difficult to reproduce if the process only follows a predefined path.
None of this means organisations need to rip out their existing security tooling or dramatically change their teams. It means changing what those investments are expected to prove. There is an important difference between saying, “we deployed the control” and saying, “we tested it against a realistic attack and know where it works and where it doesn’t.” The first is a statement; the second is evidence.
Cyber-risks are now material business risks, so security leaders need to give boards something stronger than technical activity metrics and assurances of good intent.
Absolute security is not a realistic target for an environment that changes every day. A more useful goal is to understand your most important exposures, know which controls work against them, and identify gaps while you can still fix them. An assumption can survive for years if nobody tests it, but an attacker only needs to test it once.
Alan Simpson is Field CISO at Rapid7
Main image courtesy of iStockPhoto.com and gorodenkoff
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543