AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science
Abstract
Large language models are increasingly used to support scientific validation. We experimentally compare human-only, AI-assisted, and AI-led teams assessing the reproducibility of quantitative social science research. Human-only and AI-assisted teams achieved comparable reproduction rates and substantially outperformed AI-led teams, while human-only teams detected more major coding errors. The results indicate that expert human judgment remains essential for reliable empirical verification.
Type
Publication
Proceedings of the National Academy of Sciences, 123(22), e2524747123