Agents finally got a security exam. duma-bench, out september 21, tests llm agents under dual…
agents finally got a security exam. duma-bench, out september 21, tests llm agents under dual control with eight attack classes from rag poisoning to cross-agent manipulation across 14 models.
your prompt injection surface is now a scored leaderboard.
Context
The DUMA-Bench paper on arXiv, listed on September 21, 2026 and written by authors at AI Security Lab at ITMO University and Hive Trace Lab, extends tau2-bench with adversarial environments in which both the agent and the user can change shared environment state, which it calls dual control. It describes eight vulnerability classes across eight domains, including retrieval poisoning, cross-agent manipulation, unsafe downstream output, trusted-data oversharing, identity spoofing and tool shadowing, and tests 14 models from five families. It uses 35 executable tasks, 5 runs per model-domain-task configuration, a simulated user played by GPT-4o-mini, and mostly deterministic environment assertions with LLM-judged communication assertions in 9 tasks. It reports an aggregate attack success rate of 26.9% under passive-user evaluation against 41.1% under dual control.
The results are author-reported from a preprint not known to be peer reviewed, and the authors list modest size, simulator realism and judge variance as limitations. The paper says it is intended as a methodological benchmark and not a model leaderboard, so a security exam for agents is the author's framing. Per-model results were not read. The repository, which the paper points to, was inspected and carries an MIT license.
Watch next
- Per-model results and independent replication.
Sources
- arXiv: DUMA-Bench (September 21, 2026)arxiv.org
- DUMA-Bench repositorygithub.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 10:06 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →