The agent is fastest on code you don't know. a metr randomized trial, reported september 19, found…
the agent is fastest on code you don't know. a metr randomized trial, reported september 19, found experienced developers using ai tools were 19% slower on familiar codebases, while the speedup showed up on unfamiliar work.
the tool taxes your strongest territory.
Context
METR's study of 10 July 2025 was a randomized controlled trial with 16 experienced open-source developers on their own repositories (22k plus stars, over 1M lines of code), where allowing early-2025 AI tools made them take 19 percent longer to finish issues. METR's update of 24 February 2026 says it is changing the study design, that the newer data is an unreliable signal because of selection effects, that developers are likely more sped up now, and that 30 to 50 percent of developers avoided submitting some tasks they did not want to do without AI.
The 19 percent is first-party but it is the July 2025 study on early-2025 tools, on developers' own repositories, and no METR item dated 19 September was found. A speedup on unfamiliar work was not found in the text read, so that part is unsupported. METR's February 2026 update points toward a speedup for later data while calling it unreliable, so the 19 percent is not a current finding. The tool taxes your strongest territory is the author's take.
Related work
- Earlier note on METR and the productivity story ↗Cites the same two METR posts.
- Earlier note on METR's rerun attempt ↗The 2026 update to the same study.
Watch next
- Any METR item dated September 2026 and a familiar versus unfamiliar breakdown.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 14:15 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →