AIORG-W011 Preprints

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

METR developer RCT · arXiv:2507.09089

preprint abstract and full text

Recorded claims
2
coded from the inspected source
Evidence classes
1
adverse-or-corrective
Recorded limitations
5
stated, never hidden

From the wave-0 evidence ledger

Recorded claims

Each claim states what the inspected source says, at the recorded location—bounded by its scope and graded by its confidence. Nothing here is a synthesis across works.

Across 246 real tasks performed by 16 experienced open-source developers in repositories they knew well, allowing early-2025 AI tools increased completion time by 19% on average.

confidence: high within study; medium external validity adverse-or-corrective

Scope: Experienced maintainers in mature open-source repositories; small sample and specific tooling.

Participants nevertheless predicted that AI would make them about 20% faster, showing a large perception-versus-measured-performance gap in this setting and a reason to instrument adoption rather than rely only on user belief.

confidence: high within study adverse-or-corrective

Scope: Same small expert sample; does not establish a general user-perception bias magnitude.

Boundaries

Limitations & independence

Recorded at coding time, carried with the work forever. A claim without its limits is not evidence.

Recorded limitations

  • 16 experienced developers
  • 246 tasks in mature open-source repositories
  • Developers had about five years' repository familiarity on average
  • Early-2025 tools and interfaces
  • Preprint

Source independence

Nonprofit research organization study; not authored as a software-vendor case in the inspected record.