← Founder Notes
Archive

Ai assistants can fix the bugs they cannot find. a new swe-explore benchmark, out september 17,…

Yethikrishna ROriginal on Threads

ai assistants can fix the bugs they cannot find. a new swe-explore benchmark, out september 17, shows coding assistants fall to 14-19% accuracy at line-level bug localization even when they repair whole files well.

finding the bug is now the harder half.

Context

The SWE-Explore preprint (arXiv 2606.07297) builds a benchmark of 848 issues across 10 languages and 203 repositories that isolates repository exploration: explorers return ranked code regions under a line budget, with ground truth derived from independent agent trajectories that solved each issue. The paper text says general-purpose coding agents all reach high HitFile and nDCG@500 while their line-level recall (Rec-l) stays around 0.14 to 0.19.

How it compares

The 0.14 to 0.19 range is line-level recall, not accuracy, so accuracy is the note's wording. Out September 17 is not supported: search metadata dates the paper 5 June 2026 and the repository 8 June 2026. Repair whole files well rests on strong file-level hit rates in the paper, and no measured whole-file repair result was confirmed. Results use a Mini-SWE-Agent scaffold across models. Preprint, not peer reviewed.

Watch next

  • A dated release for the benchmark and any whole-file repair measurement.

Sources

  1. SWE-Explore (arXiv 2606.07297)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 20:16 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/ai-assistants-can-fix-the-bugs-they-cannot-Ddg1VdpgJHx" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Ai assistants can fix the bugs they cannot find. a new swe-explore benchmark, out september 17,…"></iframe>

More notes