ao link
Menu
Teiss - Cracking Cyber Security
Teiss - Cracking Cyber Security

What we learnt from AI-driven vulnerability discovery

Since the launch of Mythos, numerous AI companies, chief among them Anthropic itself, have claimed that vulnerability management has fundamentally changed forever.

 

With so much hype and discussion, it can be difficult for researchers to work out how frontier AI will change their jobs. At OPSWAT, where I’m Senior Team Leader of Unit 515, its internal red team and cyber-security research unit, we’ve integrated AI into our vulnerability discovery process and, in the first few months, have publicly disclosed 18 CVEs across 12 distinct ecosystems.

 

So let me reveal how we implemented AI into our processes, how it went, and my advice for other research divisions going through the same situation.

 

How we got to AI-driven vulnerability remediation

At this stage in our vulnerability discovery programme, we’ve implemented frontier models and agentic workflows that can now discover, validate, prioritise, and remediate vulnerabilities.

 

Our objective was not to chase model hype or treat any single frontier model as a universal bug-finding engine. Instead, we wanted AI to be engineered into a practical, repeatable, and high-confidence vulnerability research pipeline that produces measurable, real-world security impact.

 

At first, we took the most straightforward approach of giving a large codebase to a frontier model and asking it to find vulnerabilities. It worked well initially. On small, isolated code samples, the model could identify suspicious patterns and even suggest plausible exploit payloads.

 

However, as we attempted to scale this approach to real-world codebases, it quickly began to fall apart. Large codebases introduce architectural complexity, implicit trust boundaries, and relationships across system architecture.

 

As the context grew, the model’s output became less reliable, producing findings that sounded technically convincing but were not actually exploitable. Some exploits depended on unrealistic assumptions, such as already having administrator-level access, while others pointed to code that was unreachable in practice.

 

In the end, the model was making the situation worse. Every AI-generated finding still had to be manually verified, wasting even more of our time while increasing AI usage costs without delivering proportional security value.

 

We soon realised that methodology was far more important than the choice of model in AI-driven vulnerability discovery. Whilst more capable models can accelerate discovery, it is human domain expertise and disciplined validation that produce consistent, high-confidence security findings.

 

Our second approach was to engineer AI into a human-led vulnerability research pipeline that could be continuously refined, validated, and improved over time.

 

Why human validation was central to everything

The challenge with instructing AI to simply go and find bugs is that models lack domain expertise, security methodology, and a disciplined workflow.

 

During our first approach, the model did not inherently understand which paths were attacker-controlled, which functions were security-critical, which assumptions were realistic, or which areas were in or out of scope. It takes security expertise to determine whether a piece of code represents a real, exploitable vulnerability.

 

So, when we reworked our approach, we divided the task into clear stages rather than asking a model to discover vulnerabilities end to end.

 

Those stages included understanding the codebase, mapping the attack surface, identifying sources and sinks, tracing data flows, analysing permission boundaries and reachability, developing exploit hypotheses, validating impact, reviewing patches, and preparing responsible disclosures.

 

By breaking the process into stages, we gave the model a more effective and controlled role. We provided each phase with a focused task, clearer context, and a specific expected outcome. This approach reduced the risk of the model drifting out of scope, losing important context, or producing hallucinated findings.

 

Human intuition identified where deeper analysis was needed, while AI helped expand the investigation and accelerate analysis.

 

Why having the best AI model isn’t the defining factor

One of the biggest lessons we learned during this process was that methodology was more important than the model.

 

As competition between AI providers intensifies, it can be tempting to buy the latest and most powerful model, but that is not always the right choice. In many cases, it simply becomes a cost multiplier. The most capable model is not required at every stage.

 

Smaller or lower-cost models may be sufficient for tasks like summarisation, classification, code navigation, or initial pattern matching. More capable models, meanwhile, should be reserved for deeper analysis and complex cross-file investigations.

 

Ultimately, process maturity matters more than raw model capability. AI can dramatically improve vulnerability research, but only when it is paired with human domain expertise, structured workflows, and evidence-based validation.

 

Toeing the line between AI-driven research and AI slop

Throughout the entire process, we were very conscious of not carelessly creating more noise. In fact, our defining motto is not to generate more reports but to generate better security outcomes.

 

Public reporting has documented how AI-generated submissions are steadily increasing the triage burden on bug bounty teams, and we did not want to become part of the problem.

 

Therefore, every submission must be grounded in evidence. That includes affected versions, technical root cause, reproducible conditions, realistic impact, and clear remediation guidance. AI can expedite the path to a hypothesis, but it does not become a vulnerability until it has gone through expert review.

 

Every AI-generated hypothesis should be tested, and if false positives continue to appear, the workflow should be improved rather than assuming that a more capable model alone will solve the problem.

 

Ultimately, during this entire process, we found that the next era of vulnerability research will not be won by teams that simply use the most capable model. It will be won by teams that build the most disciplined research systems around AI.

 


 

Loc Nguyen is Senior Penetration Test Team Lead at OPSWAT

 

Main image courtesy of iStockPhoto.com and Alexander Sikov


Please take 30 seconds to register

Register Now

 

Already have an account? Sign in

Remember Login
Teiss - Cracking Cyber Security

Subscribe to our Weekly Newsletter

Receive the latest insights direct to your inbox, and gain access to our exclusive events.
Teiss - Cracking Cyber Security

Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF

 

020 8349 4363

info@teiss.co.uk

 © 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543