Coding benchmarks overrate code. kimi k3, evaluated september 21, hits 95.10% on the swe-bench…
coding benchmarks overrate code. kimi k3, evaluated september 21, hits 95.10% on the swe-bench verified subset yet lands eighth of 43 models on the vals index with 57.81%.
the gap between a coding subset and a real task is the story.
Context
Vals AI's Kimi K3 page, with a Vals Index evaluation post dated 16 July 2026, says Kimi K3 scores 57.81 percent on the Vals Index (plus or minus 1.06), number 8 of 43 models, with component results of 95.10 percent on the SWE-bench Verified Vals Index subset, 91.27 percent on the Vibe Code Bench subset and 80.90 percent across three trials of Terminal-Bench 2.1, at temperature 1, up to 262k output tokens and max-effort reasoning. A third-party aggregator shows 57.81 as a Vals Index v2.1 entry.
The figures 95.10, 57.81 and number 8 of 43 are first-party from the evaluator, as of the page's evaluation. Evaluated September 21 is not supported, since the page dates the evaluation to 16 July 2026. It is a Vals Index subset of SWE-bench Verified and not the full benchmark. Coding benchmarks overrate code is the author's inference: the Vals Index is a composite with non-coding components such as finance tasks, so a lower composite rank does not show coding scores are overstated, and the same page reports high coding components. A ranking among 43 models at one date is not a matched coding-only comparison.
Watch next
- Vals' current Kimi K3 listing and index version.
Sources
- Kimi K3 (Vals AI)vals.ai
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 11:17 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →