
The security industry is targeting the wrong threat. The current conversation about AI-accelerated discovery fixates on capability: what these systems can find, how fast, and who gets access. That was never the step that kept organisations safe.
Discovery speed has been climbing for years, and AI is pushing it higher. Models capable of surfacing candidate exposures at scale are already in production, with broader rollout underway. But finding a candidate for exposure has never been the same as having a finding.
Offensive security operates under an evidentiary standard: a finding must be reproducible under specific conditions and demonstrably exploitable in a way that is impactful in the organisational context. As discovery gets cheaper, validation becomes the scarce resource, because it is where irreducible human judgement lives.
Teams bracing against what AI can find are bracing against the wrong threat. The ones who will struggle are those who let an AI-flagged hypothesis enter the workflow, carrying the same status as a validated finding.
"Validation" is not one thing
People use the word validation as if it describes a single gate. It describes a sequence of distinct steps, and those steps do not commoditise equally.
The evidence standard is precise: a finding is reproducible under specific conditions and demonstrably exploitable in a specific way. A confidence score from a model is a hypothesis, not a finding. That distinction matters because findings leave the security team and reach developers, clients, compliance, and leadership. Every one of those audiences asks the same question: how do you know?
Walking the post-discovery sequence honestly reveals which parts of validation are habit dressed up as expertise and which require judgment that cannot currently be delegated.
Where judgement is irreducible and where it is habit
Deduplication is largely commoditisable. Pattern matching at scale is what automated systems do well, and conceding this cleanly is more useful than defending it.
Reproduction is partly mechanical. A deterministic tool can re-run the check, but deciding which environmental conditions matter and whether the test environment reflects the production system requires judgment.
Confirming exploitability is the crux of the sequence. Reachable, triggerable, and impactful are three different questions. Compensating controls may exist; the vulnerable path may not be reachable in this deployment. Answering these requires understanding the system, not just the CVE. Deterministic exploit validation can produce the proof, but deciding what to validate and reading the result in context is a judgment that AI can assist with, but not replace.
Assessing blast radius requires business and architectural context that no model holds: what the asset touches, what sits downstream, and what a successful exploit would actually unlock.
Prioritisation uses KEV, EPSS, and CVSS as inputs, not as decisions. The right call is sometimes that a medium stays a medium, and that call is environment-specific.
Handoff, writing something a developer can act on, bridges two contexts and resists automation for that reason. A finding summary that satisfies a security analyst will not necessarily tell a developer what to change or why the change matters.
Tracking and risk acceptance are governance questions, not a queue artefact. The distinction between deliberate acceptance and informal backlog tolerance is a judgment call, and it does not get easier when the discovery engine speeds up.
The pattern holds throughout: the front of the validation sequence commoditises, the back of it does not. Being honest about which is which is itself a form of expertise.
What the data shows
In our data, across all vulnerabilities Pentest-Tools.com customers surfaced in 2026 through 28 May, approximately 25% were confirmed automatically, with proof, through active attacks and exploitation. The remaining 75% were unconfirmed and required additional validation.
Automated exploitation carried a quarter of findings end-to-end and produced reproducible evidence. Real work done. But it did not remove the validation step for the other three-quarters; it concentrated the effort where judgment is required.
One caution worth making explicit: unconfirmed means a finding that has not yet cleared the evidence bar. Both treating that 75% as immediately actionable or dismissing it as noise are approaches which skip the step that matters.
The real failure mode: badge inflation
When an AI-flagged hypothesis carries the same status as a validated finding, the result is a dashboard that cannot distinguish between assessed-and-accepted and not-yet-reviewed. This is a known failure in vulnerability management programmes generally: the inability to tell calculated risk acceptance from capacity failure dressed up as a decision. AI-accelerated discovery scales it.
A workflow that auto-confirms 25% of findings and lets the other 75% drift forward with equivalent status erodes trust in the queue. Once that trust is gone, it is hard to rebuild.
How this reshapes the team and the workflow
If validation is the scarce step, the practical implications follow directly:
Internal discipline predates AI and survives it
AI improves the top of the funnel: deduplication, coverage, triage overhead, and volume. That is genuinely useful, and teams that use it well will move faster at the front of the sequence and have more capacity for the work at the back.
Validation is where the evidence standard lives, and that standard did not move when the discovery engine got faster. Understanding that tells you where to place your people and how to build the skills they will need at the next stage of their careers.
The discovery engine got faster. The standard of proof did not move.
Daniel Bechenea is Product Security Manager at Pentest Tools
Main image courtesy of iStockPhoto.com and spawns
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543