Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
- Published
- Source
- arXiv
- Paper number
- 1164
- Field
- safety
- arXiv ID
- 2609.40286
Key points
- It built a cross-lingual unlearning benchmark spanning 174 language-script pairs and 25 paraphrase types to measure when forgetting spreads across languages.
- It showed the surprising result that collecting individually strong languages does not reliably form a good source set, and proposed coverage-prediction-based COVER.
- COVER needs only benign calibration data and access to the frozen model at deployment, reducing scarring without full multilingual retraining.
- Across three model families and two forget sets it lowered residual access by 7.8-27.3% on average versus uniform selection, and the effect held on real low-resource news documents.
Paper links
External research summaries. These are not HDATF publications or measured product results.