The code review benchmark now has a clear leader. augment's review agent, powered by gpt-5.2 and…
the code review benchmark now has a clear leader. augment's review agent, powered by gpt-5.2 and out september 17, beat cursor bugbot and coderabbit by about 10 points on the only public benchmark for ai-assisted review.
reviewing well is now a measured capability.
Context
Augment's blog on why GPT-5.2 is its model of choice for Augment Code Review, dated 11 December 2025 and last updated 18 June 2026, says it has the highest accuracy on the only public benchmark for AI-assisted code review, outperforming systems from Cursor Bugbot, CodeRabbit and others by about 10 points on overall quality. The Martian code-review-benchmark README describes an open replication of the benchmark used by Augment and Greptile: 50 pull requests across 5 codebases, human golden comments and an LLM judge, with judge variance and sparse nit coverage noted as limits.
The roughly 10 point gap is a vendor statement on one benchmark with one scoring setup; the page read gives no numbers, benchmark name or judge. Out September 17 is not supported, since the post dates to December 2025 with a June 2026 update. The only public benchmark is Augment's wording; secondary posts describe other benchmarks with other winners, which are different benchmarks and not refutation. The replication README gives no score that was read. Reviewing well is now a measured capability is the author's take.
Watch next
- Augment's full benchmark analysis and an independent run of the replication.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 21:20 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →