← Founder Notes
Archive

Benchmarks just moved from answering questions to doing ai research. vals sage, updated september…

Yethikrishna ROriginal on Threads

benchmarks just moved from answering questions to doing ai research. vals sage, updated september 10, grades 81 models on autonomous llm r&d against published, human and frontier references.

the exam is now can a model build the next model.

Context

Vals AI's benchmark listing has two entries relevant to this note. The page named SAGE is Student Assessment with Generative Evaluation, which grades handwritten math work, was updated September 1, 2026 and shows 23 models, with Claude Opus 4.7 leading at 56.10 percent. A separate Vals RSI Index page, updated August 12, 2026 with 5 tasks and 3 named model harnesses, asks whether a model can do the research that builds the next model, and the benchmarks listing describes it as autonomous LLM R&D scored against published, human and frontier-model references.

How it compares

The note names SAGE but its description matches the RSI Index listing, and which one the author meant is not asserted here. The 81 models and the September 10 date were not found on either page read, so those two figures are unsupported, not refuted. Scores must not be merged across Vals benchmarks or across harnesses. The exam is now can a model build the next model is the author's take.

Watch next

  • Whether the RSI Index updates with a larger model set, and the source of the 81 models.

Sources

  1. Vals AI: SAGEvals.ai
  2. Vals AI: RSI Indexvals.ai
  3. Vals AI: benchmarksvals.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 14:37 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/benchmarks-just-moved-from-answering-questions-to-doing-DdlYEnFjXv7" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Benchmarks just moved from answering questions to doing ai research. vals sage, updated september…"></iframe>

More notes