Introducing Nested Learning: A new ML paradigm for continual learning

Admin

Introducing Nested Learning: A new ML paradigm for continual learning

A paper published at NeurIPS 2025 introduces Nested Learning, a machine learning approach aimed at addressing continual learning and catastrophic forgetting in large language models (LLMs).

The source text says recent progress in machine learning has been driven by powerful neural network architectures and the algorithms used to train them, but that important challenges remain around continual learning, where a model acquires new knowledge and skills over time without forgetting old ones.

According to the paper, titled “Nested Learning: The Illusion of Deep Learning Architectures”, existing LLMs are limited to either the immediate context of their input window or the static information learned during pre-training. The text compares this to the human brain’s neuroplasticity and says current models lack a similar ability to adapt continuously.

The paper argues that continually updating a model’s parameters with new data can lead to “catastrophic forgetting”, where learning new tasks reduces performance on earlier tasks. It says researchers typically try to reduce this through architectural changes or improved optimization rules, but that these have often been treated as separate parts of the system.

Nested Learning is described as a way to bridge that gap by treating a single ML model as a system of interconnected, multi-level learning problems optimized simultaneously. The paper says the model’s architecture and the rules used to train it are fundamentally the same concepts, operating at different levels of optimization with different “context flow” and update rates.

The authors say this structure creates a new dimension for designing AI systems and can help build learning components with deeper computational depth, which they say may help address catastrophic forgetting.

The paper also presents a proof-of-concept self-modifying architecture called “Hope”. The source says it achieves superior performance in language modeling and better long-context memory management than existing state-of-the-art models.

Source: research.google.

Companies can share verified announcements through Newz9’s international press release submission page.