Penetration Tester.
AI agents can now find real vulnerabilities at scale, compressing routine testing work while pushing human testers toward scoping, exploit chaining, and business judgment.
Medium risk, high transformation.
Adoption of AI testing tools is high and rising fast, and the role sits in the heavily transformed part of the map. Penetration testing is not disappearing: demand for security testing keeps growing, and the federal government projects 29% employment growth through 2034 for information security analysts, the broader occupation that includes penetration testers. What is shifting is the daily work. The routine discovery tasks at the core of the job are being automated fastest, so the human value is moving up toward scoping, exploit chaining, judgment about business impact, and accountability for results.
3 shifts already visible in the data, in order of magnitude.
An AI system reached the top of HackerOne's US bug-bounty leaderboard.
In June 2025, XBOW became the first non-human to top HackerOne''s US leaderboard, submitting nearly 1,060 vulnerability reports in a few months. On an internal benchmark of 104 challenges it matched a veteran pentester''s 40-hour result in 28 minutes, a direct signal that the routine, high-volume end of testing can now be automated.
AI agents are finding real zero-days in production code.
Google''s Big Sleep agent, built by DeepMind and Project Zero, identified a critical SQLite zero-day that was known only to threat actors, before it could be exploited, and has reported roughly twenty previously unknown flaws in widely used open-source software.
DARPA's AI Cyber Challenge produced systems that find and patch bugs on their own.
The two-year competition concluded at DEF CON 2025 with AI systems that autonomously detect, exploit, and patch vulnerabilities in open-source software that underpins critical infrastructure. Team Atlanta''s ATLANTIS won the $4M grand prize, and all seven finalist teams are open-sourcing their systems.
What the leaders are doing.
| № | Company | Sector | What they are doing | Year | Source |
|---|---|---|---|---|---|
| 01 | XBOW | Security | Its autonomous AI penetration tester became the first non-human to top HackerOne's US bug-bounty leaderboard, submitting nearly 1,060 vulnerability reports in a few months and matching a veteran pentester's 40-hour benchmark in 28 minutes. | 2025 | darkreading.com ↗ |
| 02 | Technology | Its Big Sleep agent, built by DeepMind and Project Zero, found CVE-2025-6965, a critical SQLite zero-day known only to threat actors, before it could be exploited, and has reported roughly twenty previously unknown flaws in open-source software. | 2025 | therecord.media ↗ | |
| 03 | DARPA | Government | Ran the two-year AI Cyber Challenge, concluded at DEF CON 2025, where AI systems autonomously found and patched vulnerabilities in open-source software behind critical infrastructure. Team Atlanta's ATLANTIS won the $4M grand prize; all seven finalist teams are open-sourcing their tools. | 2025 | cybersecuritydive.com ↗ |
What is declining, growing, emerging.
- 01Routine web-application and external network scanning on well-understood vulnerability classes
- 02First-pass reconnaissance and enumeration done by hand
- 03High-volume, low-complexity bug bounty submissions on common flaw types
- 01Scoping engagements and defining rules of engagement for a specific business and threat model
- 02Chaining multiple low-severity findings into a demonstrable, high-impact attack path
- 03Judging real-world exploitability and business impact given a client's environment and controls
- 04Reviewing and validating findings produced by autonomous testing tools, and filtering false positives
- 01Testing AI systems themselves: prompt injection, agent permission abuse, and model supply-chain weaknesses
- 02Operating and directing autonomous pentest agents as a force multiplier across a wider attack surface
- 03Adversarial red teaming of AI-powered defenses and detection systems