← Founder Notes
Archive

Just shipped a build with agents fully banned from touching the test file. same task, 9 hours, 2.1M…

Yethikrishna ROriginal on Threads

just shipped a build with agents fully banned from touching the test file. same task, 9 hours, 2.1M tokens, cost dropped 38% and the tests actually pass on first run now.

ryanvogel was right.

Context

Anthropic's post on building effective agents covers testing and evaluation inside agent loops and recommends simple, composable patterns. METR's early-2025 study found experienced open-source developers took 19 percent longer with AI tools in that setting.

How it compares

The build, the 9 hours, 2.1M tokens and 38 percent cost drop are the author's own measurements, and I inspected no evidence for them. Neither source above says anything about banning agents from test files, and the ryanvogel reference is not identified.

Watch next

  • Whoever ryanvogel refers to, if the author names the source.

Sources

  1. Building effective agents (Anthropic)anthropic.com
  2. Early-2025 AI and experienced open-source developer study (METR)metr.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 3 September 2026 at 21:01 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/just-shipped-a-build-with-agents-fully-banned-Dc1I6cKjapv" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Just shipped a build with agents fully banned from touching the test file. same task, 9 hours, 2.1M…"></iframe>

More notes