Alibaba's open code review beats claude code on its own benchmark with a ninth of the tokens, then…
alibaba's open code review beats claude code on its own benchmark with a ninth of the tokens, then an independent run scored about 12 percent precision.
the gap between vendor benchmarks and outside verification is the real story.
Context
The alibaba/open-code-review README says that compared with general-purpose agents (Claude Code) it achieves significantly higher precision and F1 with the same underlying model while using about one ninth of the tokens, and that its recall is lower. The benchmark covers 50 open-source repositories, 200 real pull requests, 10 languages, more than 80 senior engineers and 1,505 annotated issues. A Hacker News commenter reported running the tool on 10 of the 50 pull requests of a different benchmark, a Martian code review benchmark, and getting recall about 74 percent, precision about 12 percent and F1 about 20 percent.
Beats Claude Code holds for precision and F1 only, since the README says recall is lower, and it is vendor-reported on the vendor's own benchmark. The independent run is one unverified commenter's report on a subset of a different benchmark with a different judge and setup, so it is a separate result and not a head-to-head test; no matched Claude Code comparator exists, so the 74 percent recall neither supports nor conflicts with the README. The two sets of numbers are not merged.
Related work
- Earlier note on Alibaba's open-code-review ↗Same repository and README claims.
Watch next
- Independent runs on the full 50 pull request set and the paper's numeric tables.
Sources
- open-code-review README (Alibaba, GitHub)github.com
- Hacker News comment on running open-code-reviewnews.ycombinator.com
- OpenCodeReview (arXiv 2608.09290)arxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 10:38 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →