Unlearning is now a hallucination fix. llm ghostbusters, updated september 17, cuts package…
unlearning is now a hallucination fix. llm ghostbusters, updated september 17, cuts package hallucination rates by 81% by surgically unlearning bad outputs while keeping the model's general ability.
the fix is forgetting, not remembering.
Context
The arXiv preprint 2605.01047 proposes Adaptive Unlearning, a closed-loop framework with a token-level objective that reinforces valid and suppresses hallucinated package names, using synthetic data from the model itself with no human labels. Version 1 (1 May 2026) says it reduces package hallucination rates by 81 percent. Version 2, now titled Surgical Package Hallucination Suppression via Adaptive Unlearning, says 88 percent while maintaining performance on standard coding benchmarks, with rates of 2.56 percent on DeepSeek-Coder-7B (from 21.23) and 6.13 percent on DeepSeek-Coder-V2-Lite (from 27.75), on held-out prompts, with utility checked on EvalPlus pass@1.
The note's 81 percent is the version 1 figure and the later version states 88 percent. An updated September 17 date was not found, since search metadata dates version 2 to 21 September. The scope is two models of one family and one hallucination type, package names, not hallucination in general, and keeping general ability rests on EvalPlus pass@1 only. It is a preprint and its code link was not opened. Surgically is the note's word.
Related work
- Earlier note on slopsquatting ↗Package hallucination as an attack surface, a different measure from this mitigation.
Watch next
- The peer-reviewed version and evaluation on other model families.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 17:48 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →