Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

Published
Source
arXiv
Paper number
1164
Field
safety
arXiv ID
2609.40286

Key points

  • It built a cross-lingual unlearning benchmark spanning 174 language-script pairs and 25 paraphrase types to measure when forgetting spreads across languages.
  • It showed the surprising result that collecting individually strong languages does not reliably form a good source set, and proposed coverage-prediction-based COVER.
  • COVER needs only benign calibration data and access to the frozen model at deployment, reducing scarring without full multilingual retraining.
  • Across three model families and two forget sets it lowered residual access by 7.8-27.3% on average versus uniform selection, and the effect held on real low-resource news documents.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)