← Founder Notes
Archive

Alibaba's open code review beats claude code on its own benchmark with a ninth of the tokens, then…

Yethikrishna ROriginal on Threads

alibaba's open code review beats claude code on its own benchmark with a ninth of the tokens, then an independent run scored about 12 percent precision.

the gap between vendor benchmarks and outside verification is the real story.

Context

The alibaba/open-code-review README says that compared with general-purpose agents (Claude Code) it achieves significantly higher precision and F1 with the same underlying model while using about one ninth of the tokens, and that its recall is lower. The benchmark covers 50 open-source repositories, 200 real pull requests, 10 languages, more than 80 senior engineers and 1,505 annotated issues. A Hacker News commenter reported running the tool on 10 of the 50 pull requests of a different benchmark, a Martian code review benchmark, and getting recall about 74 percent, precision about 12 percent and F1 about 20 percent.

How it compares

Beats Claude Code holds for precision and F1 only, since the README says recall is lower, and it is vendor-reported on the vendor's own benchmark. The independent run is one unverified commenter's report on a subset of a different benchmark with a different judge and setup, so it is a separate result and not a head-to-head test; no matched Claude Code comparator exists, so the 74 percent recall neither supports nor conflicts with the README. The two sets of numbers are not merged.

Related work

Watch next

  • Independent runs on the full 50 pull request set and the paper's numeric tables.

Sources

  1. open-code-review README (Alibaba, GitHub)github.com
  2. Hacker News comment on running open-code-reviewnews.ycombinator.com
  3. OpenCodeReview (arXiv 2608.09290)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 10:38 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/alibaba-s-open-code-review-beats-claude-code-DdiX8xoClND" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Alibaba's open code review beats claude code on its own benchmark with a ninth of the tokens, then…"></iframe>

More notes