← Founder Notes
Archive

Benchmarks finally stopped grading diffs. backendforge, out since july, runs 56 backend tasks where…

Yethikrishna ROriginal on Threads

benchmarks finally stopped grading diffs. backendforge, out since july, runs 56 backend tasks where an llm must ship a dockerized service that passes http tests against an openapi contract.

agents pass or fail on shipped behavior.

Context

BackendForge is described in an arXiv paper listed on July 13, 2026 by authors at Peking University and Fudan University. It has 56 contract-defined backend generation tasks rewritten from real open-source applications, with OpenAPI contracts defining 2,345 API operations. Given a visible spec and contract, a model must produce a Dockerized service that is built, deployed and evaluated only through HTTP tests, and scoring does not inspect source code. A base oracle is made first, then a test agent and a code agent co-evolve the test oracle and reference service, with a review agent filtering tests. The authors report the best model solving 55.4% under the base oracle and 28.6% under the final oracle (16 of 56 tasks), and every other configuration at most 3 of 56 under the final oracle.

How it compares

The results are the authors' own from a preprint, and the base oracle and final oracle are different measures. The paper's limits are Python backend services tested through black-box HTTP only, with no UI, other stacks, performance, deployment security or production hardening. It positions itself against repository-repair benchmarks and itself cites earlier backend benchmarks that grade deployed behavior, BaxBench and ABC-Bench, so finally is the author's framing.

Watch next

  • Replication on other stacks.

Sources

  1. arXiv: BackendForge (July 13, 2026)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 12:19 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/benchmarks-finally-stopped-grading-diffs-backendforge-out-since-DdlIQwYjZqp" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Benchmarks finally stopped grading diffs. backendforge, out since july, runs 56 backend tasks where…"></iframe>

More notes