The smartest ai code review tool this month is the one that barely uses ai. alibaba's…
the smartest ai code review tool this month is the one that barely uses ai. alibaba's opencodereview, open-sourced september 20 under apache-2.0, handles file selection and grouping deterministically and only calls the model for what a reviewer actually needs.
the ai review delusion is dying in plain sight.
Context
The OpenCodeReview paper (arXiv 2608.09290, Alibaba Group, Nanjing University and Peking University authors) says its Rule-Guided Dispatch uses a multi-layer rule system to deterministically select files and review criteria, then file-level parallel SubAgents review with a curated tool set, and an Independent Reflection step filters hallucinated comments. The authors report that on AACR-Bench, 200 real-world pull requests in 10 languages with 1,505 expert-verified comments, it outperforms Claude Code and Codex across six LLM backends, up to 2.17x higher SEM-F1 (25.10% vs 11.57% for the same model under Claude Code) with 5 to 15 times fewer tokens.
These results are the authors' own, on one benchmark, and not independent. The tool still calls an LLM for per-file review, so grouping and calling the model only for what a reviewer needs is the note's paraphrase. The paper's ID dates to August 2026, and an earlier inspection of the repository recorded a creation date of May 18, 2026 with Apache-2.0, which matches the license. A creation date is not necessarily the open-sourcing date, and no first-party September 20 source was read, so that date is not verified, not disproven. Smartest is subjective, as the paper claims better SEM-F1 and lower token use only on AACR-Bench. A separate independent run on a different benchmark, noted in the linked entry, is not a matched refutation.
Related work
- Alibaba's open code review, half rules and half LLM ↗The earlier note on the hybrid design and its trending run.
- Open Code Review beats Claude Code on its own benchmark ↗The earlier note on the vendor benchmark claim and an independent run that scored about 12 percent precision on a different benchmark.
Watch next
- Independent runs on a matched benchmark, and the repo's first public release tag.
Sources
- arXiv 2608.09290: OpenCodeReviewarxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 22:16 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →