The important change: AI has to prove the bug

Security reports can sound convincing without demonstrating a real exploit. MobileCybench approaches that problem differently. Researchers created executable probes that check whether a security property was actually violated after an agent's attempted exploit was replayed.

Five coding-agent configurations were evaluated across four settings: a malicious application on the victim's device or a remote attacker with a low-privilege account, with either an obfuscated APK or application source code available.

Top-agent results with only the obfuscated APK

A triggered probe demonstrates a checked security property failed. It is not the same as saying the agent compromised 53.8% of arbitrary Android apps.

Malicious app
53.8%
Remote attacker
16.7%

Source: Zhang et al., MobileCybench, submitted September 21, 2026. Results are pass@2 benchmark trigger rates for the top agent in these specific settings.

What we know

Current coding agents can autonomously perform meaningful parts of vulnerability discovery. In MobileCybench, the strongest tested agent triggered executable security probes in multiple real applications even when given only an obfuscated APK. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, with the paper reporting that a majority had been confirmed by maintainers.

What we don't know

This benchmark does not establish that AI can reliably compromise arbitrary real-world software. It covers 13 Android applications, specific attacker privileges, five coding-agent configurations, and a finite inventory of security properties. A probe that does not trigger does not prove an application is secure, and a probe trigger does not by itself tell us the severity or real-world impact of a vulnerability.

What is still speculation

The larger societal effect depends on deployment. If vulnerability discovery becomes dramatically cheaper, defenders could continuously scan and patch software. Attackers could use similar capability to search many more targets. It is too early to know which side gains more.

The question is shifting.

We used to ask whether AI could find real vulnerabilities. Increasingly, the more useful question is whether defenders can discover, verify, prioritize, and patch them faster than attackers can exploit them.

How to prepare

Developers

Treat AI security testing as an additional layer, not a replacement for review, threat modeling, testing, and responsible disclosure.

Security teams

Prepare for higher report volume. Reproducible evidence and automated validation may become as important as finding the vulnerability itself.

Everyone else

Don't translate benchmark percentages into claims that AI can hack half of all apps. The capability is meaningful; the scope is specific.

Primary source

← Back to AI & Humanity