Ai code review became a benchmark race this week. augment's reviewer, out september 17, runs on…
ai code review became a benchmark race this week. augment's reviewer, out september 17, runs on gpt-5.2 and claims a 10 point lead over cursor bugbot and coderabbit on the public quality benchmark.
review accuracy is now a marketing number.
Context
Augment Code's blog post on why GPT-5.2 is its model of choice for Augment Code Review, dated December 11, 2025 and last updated June 18, 2026, claims the highest accuracy on the only public benchmark for AI-assisted code review, outperforming systems from Cursor Bugbot, CodeRabbit and others by about 10 points on overall quality.
The 10 point lead and the GPT-5.2 choice are Augment's own claim and not independent, from a post first dated December 11, 2025, so out september 17 is not supported. The benchmark's name, scoring and whether competitors were run by Augment were not inspected in this pass, and the matched benchmark was not freshly checked. A secondary overview seen as a snippet lists five organizations publishing code review benchmarks, which is the only basis for the note's race framing. Review accuracy is now a marketing number is the author's take.
Watch next
- Independent reruns of the benchmark.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 21:36 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →