
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
If you're building AI agents and relying on humans to catch dangerous commands, this data should worry you. Across 40,000 game runs simulating a human-in-the-loop for an AI coding agent, players missed 1 in 3 malicious commands — and this was a game where ~34% of commands were threats, far higher than real-world rates that would induce complacency. The data makes a strong case that 'human approval' is not the security guarantee most teams assume it is, especially under time pressure and alert fatigue.
Takeaways3
- Human reviewers missed 33% of threats on average, and only 20.8% of players caught all threats without over-blocking safe commands.
- Time pressure and high command volume are the enemy of meaningful human oversight — the 'human-in-the-loop' is a much weaker control than it appears.
- 7% of players approved every single command, suggesting a non-trivial portion of real users will rubber-stamp agent actions entirely.










