An Empirical Study on the Effectiveness of Security Code Review
https://people.eecs.berkeley.edu/~daw/papers/coderev-essos13.pdf
Method
"We hired 30 developers to conduct a manual code review of a small web application."
Population
"oDesk Population. As stated previously, we hired developers through the oDesk outsourcing website. This limits our population to registered oDesk users, as opposed to the population of all web developers or security reviewers."
What it does not show
A roughly 3,500-line PHP application with artificially injected flaws, which the authors note creates an artificially flawed codebase — the distribution may not match naturally occurring vulnerabilities in large systems. Measures solo review under a fixed 12-hour budget, not team review in situ, and freelance reviewers may not represent an organisation’s own staff or dedicated security reviewers.
Anne Edmundson, Brian Holtkamp, Emanuel Rivera, Matthew Finifter, Adrian Mettler, David Wagner
No reviewer found all confirmed vulnerabilities. The average found was 2.33 with a standard deviation of 1.67, about 20% found none at all, and only 17% found the missing cross-site request forgery protection. False-positive rates were bimodal, and more experience did not reliably mean more accurate or effective.
Tier I: Controlled experiment: one codebase with known seeded vulnerabilities, 30 independently hired reviewers, fixed conditions, per-reviewer outcome measurement.