qualitylab

the station

A Few Billion Lines of Code Later: Using Static Analysis to Find Bugs in the Real World

tier III/2010/Communications of the ACM 53(2)

https://web.stanford.edu/~engler/BLOC-coverity.pdf

Method

"During a checking run, error messages are put in a database for subsequent triaging, where users label them as true errors or false positives."

Population

"As of this writing (December 2009), approximately 700 customers have licensed the Coverity Static Analysis product, with somewhat more than a billion lines of code among them."

What it does not show

Self-flagged as a single, non-generalisable vendor data point. No controlled measurement of defect-finding effectiveness against other methods, and no denominator behind the 30% — it is an experiential estimate, not a measured rate on a defined sample.

Al Bessey, Ken Block, Ben Chelf, Andy Chou, Bryan Fulton, Seth Hallem, Charles Henri-Gros, Asya Kamsky, Scott McPeak, Dawson Engler

The false-positive rate is the lever that decides whether anyone acts on the output: “False positives do matter. In our experience, more than 30% easily cause problems. People ignore the tool. True bugs get lost in the false. A vicious cycle starts where low trust causes complex bugs to be labeled false positives, leading to yet lower trust.”

Tier III: Single-vendor experience report with aggregate usage numbers but largely qualitative lessons; no comparison group. The authors flag the single-data-point problem themselves.

Cited by