Continual Learning & Memory · 2025

Continual Learning via Sparse Memory Finetuning

Jessy Lin, Luke Zettlemoyer, Gargi Ghosh, Wen-Tau Yih, Aram Markosyan, Vincent-Pierre Berges, Barlas Oğuz

Demonstrated that updating only memory slots highly activated by new knowledge relative to pretraining data enables learning without catastrophic forgetting in language models.

Editorial record

Plain-language summary

Sparse memory finetuning leverages memory layer models to update only the top-ranked memory slots based on TF-IDF scores computed against background pretraining data, reducing interference between new and existing knowledge. On fact learning and document QA tasks, this approach learns new knowledge to the same degree as full finetuning while showing substantially less forgetting—NaturalQuestions F1 drops only 11% with sparse memory finetuning versus 89% with full finetuning and 71% with LoRA.

Source record

Provenance

Record ID
P-667
Record created
2026-08-12
Last reviewed
2026-08-12
Record version
1

Citation caveat: Citation metadata is approximate and marked unverified in the source dataset.