In this paper, the authors tackle continued pretraining without the risk of catastrophic forgetting, by identifying parameters which can safely be changed without risking identified concepts, and freezing the rest: https://arxiv.org/abs/2604.19089v1 Current practice is to mix new datasets into compr