qualitylab

the station

Onboarding vs. Diversity, Productivity and Quality — Empirical Study of the OpenStack Ecosystem

tier II/2021/ICSE 2021 (IEEE/ACM)/partly self-reported

https://doi.org/10.1109/ICSE43902.2021.00097

Method

"we first carry out an observational study of 72 new contributors during an OpenStack onboarding event to provide a catalog of teaching content, teaching strategies, onboarding challenges, and expected benefits. Next, we empirically validate the extent to which diversity, productivity, and quality benefits are achieved by mining code changes, reviews, and contributors' issues with(out) OpenStack onboarding experience."

Population

"Finally, we mapped the 1,281 (3x427) selected contributors across all three categories to their activities in the following OpenStack repositories: Gerrit (code review system), git, and Launchpad/Storyboard (issues trackers)."

What it does not show

Correlational throughout; the authors never claim causation. Contributors self-selected into the programme, so the groups differ in ways the matched sampling cannot remove — motivation being the obvious one. One ecosystem, seven releases, and an event run with sponsorship and dedicated mentors, so the resourcing may not transfer. Critically for this catalog: the authors state the programme was “more than giving a tutorial on creating a feature branch or running a test suite” — it measures human mentoring, not build or environment automation, and cannot be cited for the latter.

Armstrong Foundjem, Ellis E. Eghan, Bram Adams

Time to first merged commit was a median 45% lower for female and 35% lower for male and non-binary onboarded contributors, with a significant difference and large effect size against the non-onboarded group. The median probability of a commit introducing a bug was 25% without onboarding and 14% with it, at p = 4.290e-57 with a large effect size — and the onboarded group submitted more, and more complex, changes. Patch acceptance and retention also improved.

Tier II: Observational study mining 84 months of real repository history with a matched comparison group of contributors who did not onboard, plus an observational study of an event. Treatment is self-selected rather than assigned, and the authors describe every result as a correlation, so it does not meet the controlled-comparison bar for tier I. Gender, one of the diversity outcomes, is self-declared; the productivity and quality outcomes are mined from repositories.

Cited by