Skip to content

Evidence refresh: METR 2026 follow-up supersedes the 19% headline; DORA 2025 replaces 2024 as the delivery baseline #347

Description

@richardsongunde

Why this matters

RStack's research substrate (research/) and paper claims cite the METR early-2025 RCT and DORA 2024. Both have been superseded, and citing the old versions unqualified is now a claims-discipline violation of our own register (research/productivity-claims.md).

Verified in the #79 July-2026 research refresh (25/25 claims confirmed against primary sources):

  1. METR follow-up (Feb 24, 2026)https://metr.org/blog/2026-02-24-uplift-update/ — with newer agentic tools (Claude Code, Codex era), the measured effect is statistically indistinguishable from zero: returning developers −18% (CI −38%…+9%), newly recruited −4% (CI −15%…+9%). METR itself flags the data as an unreliable signal (selection bias) and is redesigning the experiment. The original page now carries an out-of-date banner. The durable finding is the perception gap (developers believed +20% while measured slower), not the 19% number.
  2. DORA 2025 (Sept 2025, ~5,000 respondents)https://dora.dev/dora-report-2025/ — AI adoption now has a positive relationship with throughput (a reversal of 2024) but still negative with delivery stability; 90% use AI, >80% self-report productivity gains. This is the strongest current empirical support for RStack's thesis: throughput-without-stability is exactly the failure mode a governed loop targets. Correlational, not causal.

Scope

  • research/bibliography.md: update the METR entry (add the 2026 follow-up, scope the 19% figure to early-2025 tools, note METR's own reliability warnings); add DORA 2025 alongside/above 2024.
  • research/productivity-claims.md: refresh the METR external-evidence row; add a DORA-2025 row (throughput/stability split); add an explicit "never cite 19% unqualified — pair the 2025 RCT with the 2026 update" rule; add the perception-gap claim as the durable citation.
  • research/methodology.md: update the reference list and the allowed-wording example that leans on METR.
  • research/paper-outline.md: "METR productivity caution" line updated to the perception-gap + stability framing.
  • Paper master (separate repo, sdlc-rstack-paper): same corrections in paper.md/paper.tex/references.bib.

Guardrails

  • Do NOT overclaim the null: METR characterizes both estimates as likely lower bounds; "no measured effect" must not become "AI does not help."
  • DORA findings are survey correlations — keep causal language out.

Source

#79 epic research refresh comment: #79 (comment)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions