What Should World Models Forget? Stratified Retention for Continual Adaptation
Continual learning treats any drop on previously seen data as failure. This paper argues that rule breaks for world models, because their prediction target is an environment that changes. Knowledge that was right when learned can become false, and dropping it is correct behavior. The authors call for retention stratified by invariance timescale. Invariants such as physics and object permanence must never be revised. Instance-level facts should be revised as soon as the world changes. Standard forgetting metrics cannot tell a good revision from catastrophic forgetting, so they rank a frozen model highest. The paper proposes differential retention: report invariant regression tests and revision latency together across the adaptation stream, without aggregation.
This paper from arXiv (2610.03713) asks a small question with a large effect: what should a continually adapting world model forget? The authors argue that the evaluation habits of continual learning do not hold for world models. One note on scope before we start. This report is based on the abstract only. The abstract reports no numerical results, so this article invents none. It analyzes the concepts, the mechanism and the evaluation claim that the paper puts forward. 1. Where the problem comes from In classic continual learning, the prediction target is stationary. A picture of a cat is a cat today and a cat in ten years. Under that condition, a drop in accuracy on an old task is a clear signal of failure. The field calls it catastrophic forgetting. Almost every forgetting metric, such as average forgetting or backward transfer, rests on this premise: the closer the old performance stays to its original level, the better the model. World models are different. Their target is the environment, and the environment changes. A warehouse moves its shelves. A junction gets a new traffic light. A robot gripper wears down and its dynamics shift. What the model learned earlier, "the shelf is on the left", was right at the time and is wrong now. A model that keeps the old answer has the real defect. In this setting, forgetting is required behavior, not a flaw. 2. The core claim: retention stratified by invariance timescale The authors say that knowledge inside a world model is not uniform. One class must never be revised. Examples are the laws of physics, the direction of gravity and object permanence, which says that an object hidden from view has not vanished. Another class is instance-level fact. Examples are the layout of one room or the present state of one machine. These should be rewritten soon after the environment changes. Between the two sits a middle band of knowledge that changes at different speeds.
This is what "stratified retention" means. The retention policy depends on the invariance timescale of each layer of knowledge. The closer knowledge sits to a true invariant, the stricter the retention requirement, and revision must be blocked. The closer it sits to a concrete instance, the looser the requirement, and revision should be encouraged and fast. The paper notes that neighboring fields already study non-stationary truth. The concept drift literature looks at shifting data distributions. Work on the temporal factuality of language models looks at facts that expire, such as who holds an office. But nobody has formulated the problem for world models. World models have one distinctive property: they also encode knowledge that must never be revised. A language model can let many facts go stale. A world model cannot let its physical common sense decay.
3. Why current metrics fail The abstract makes a sharp point. Standard forgetting metrics cannot separate two very different models. One has correctly revised outdated knowledge. The other has suffered catastrophic forgetting. Both lose accuracy on old data, so the metric sees the same signal. Worse, the metric ranks a fully frozen model highest, because a frozen model does not change at all on old data. A model that never adapts to the environment wins the contest. That defeats the purpose of evaluation. Existing physical-reasoning benchmarks have a related gap. The abstract says they evaluate only frozen checkpoints. They look at the model at one moment, not at its behavior along a continuing adaptation stream. So they do not measure continual adaptation itself. 4. Differential retention: two measures, no aggregation The proposed evaluation is called differential retention. It asks for two reports across the whole adaptation stream, side by side. First, invariant regression testing. At every adaptation stage, a test set aimed at invariants such as physics and object permanence checks that nothing was broken. This covers the rule "what must be remembered stays remembered".
Second, revision latency. After the environment changes, how long does the model take to update the affected instance-level facts? This covers the rule "what must be forgotten is forgotten quickly". The key phrase is "without aggregation". If the two are merged into one score, a model can buy a high mark on one axis by giving up the other. A frozen model scores full marks on invariants and has unbounded revision latency. Reporting them apart exposes that trade. This fits a wider caution in evaluation research against single summary scores. 5. Impact for developers and enterprises For teams that build robotics, driving, digital twins or agent simulators, the framework gives a usable checklist. First, write down an explicit set of invariants and turn it into a regression suite that runs on every update. Second, design experiments that inject controlled environment changes, and measure revision latency. Third, when you pick a continual learning method, do not look only at retention on old tasks. A method that locks old knowledge with strong regularization may look good on old data, yet it will show high revision latency under this protocol.
For benchmark builders the shift is larger. The object under test moves from a static checkpoint to an adaptation stream, the whole path of a model that is updated again and again. That raises cost. You must build sequences of environments with controlled changes, and you must label which knowledge should be rewritten and which should not. 6. Limits and open questions First, the strata are hard to draw. Which knowledge is a true invariant, and which is a fact that is merely stable over a long time? The boundary is not always clear. Physics looks eternal, but what a model holds is an approximation of physics, and an approximation may need correction. Second, measuring revision latency needs a clear label for when the environment changed. In real deployments, change is often gradual and hidden, and nobody tells the model what happened. Third, the abstract states a conceptual frame and an evaluation claim. It gives no numbers. The value of a position paper like this is that it redefines the problem. Its practical effect still has to be shown on concrete world models and concrete environments. Readers should keep this in mind when they cite it.
7. Where this may go Likely directions include benchmarks with labeled change events for world models. Another is architectures that store invariants apart from instance knowledge, for example in separate parameter subspaces or memory modules. A third is to use revision latency as a training objective. Further out, a model will need to judge for itself which knowledge has expired, which means detecting that the environment has changed. This links naturally to concept drift detection. In short, the paper proposes no new network. It redefines the problem and the scoring rule. It reminds us that when we grade a model that updates itself, we should stop asking how much it forgot. We should ask whether it forgot what it should forget, and still remembers what it must remember.