WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Published · AI Daily — AI-assisted deep research, methodology & disclosure

WOVEN offers 36,076 visual-transition examples across 20 scenes, 5 actions and 8 reasoning types. All 38 tested MLLMs trail humans. Training on subsets of about 2,000 items lifts 22 of 26 external benchmarks by up to 27.3 points.

Multimodal large language models have advanced quickly on image captioning, document understanding and visual question answering. Yet as soon as a question depends on spatial relations, embodied interaction, physical regularities or the passage of time, their performance falls behind. What happens when a robot pushes a cup toward the edge of a table? What happens to a block tower when the bottom block is pulled out? Where will an object in the frame be one second from now? These questions look different, and for years they have been treated as separate weaknesses, each measured by its own benchmark and each patched with its own data. The paper WOVEN, posted to arXiv on 8 October 2026 as 2610.12417 by Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang and colleagues, proposes a simpler reading: these failures share one root cause, a deficit in visual transition reasoning, meaning the ability to simulate internally how a scene will look after an action is applied to it.

The value of this hypothesis lies less in naming a new capability than in changing how the problem is organized. If failures in spatial, embodied, physical and temporal reasoning really do share a source, then teams need not collect expensive labels for every downstream task. Visual transition reasoning can instead serve as a shared training primitive, one that different models can learn from different supervision sources and then reuse across different tasks. The authors set out to test exactly this claim and to distill a systematic training recipe from the result. They point out that existing benchmarks document the deficits separately and do not allow controlled comparisons across scenes, actions and reasoning operations. Evaluation alone is therefore not enough. The field needs a data source organized along those three axes.

WOVEN is that data source, and it doubles as a benchmark. It contains 36,076 examples spanning 20 scene types, 5 action types and 8 reasoning types. The material comes from diverse, realistic rollouts of video-pretrained generative models: a generative model is given an initial frame and an action, it unrolls the consequences, and questions are then built around the resulting frames. This structure has a practical benefit. Researchers can hold two axes fixed and vary the third, which lets them ask questions that were hard to answer before, such as whether the scene, the action or the reasoning operation itself determines how well training transfers. Separating the variables cleanly is, in our view, the most instructive methodological choice in the paper.

The evaluation results reveal a deficit that is substantial, systematic and stubborn. The authors tested 38 frontier multimodal models, including GPT-5.4 and Qwen3-VL-235B-A22B. Even the strongest models fall far below human performance. The same failures recur across model families, so they are not quirks of one training pipeline. And the failures persist with scale, which matters most: it suggests that adding parameters and generic data does not by itself produce the ability to reason about changes in visual state. This agrees with a growing body of work on physical common sense and spatial reasoning, and it supports the case for a dedicated supervision signal.

The training experiments deliver the more useful positive result. The authors trained multimodal models at several scales on WOVEN and found that they learn a shared capability that transfers broadly. Training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks, by up to 27.3 percentage points. A second number matters even more to engineering teams: WOVEN data can replace 30 to 50 percent of a task's own training data with comparable accuracy. For embodied or spatial tasks where labels are expensive, part of the supervision could therefore come from general transition-reasoning data. A caution is in order. These figures are self-reported by the authors, four of the 26 benchmarks did not improve, and any team should retest on its own tasks before relying on them.

Finally, the paper distills a training recipe that the authors validated prospectively on held-out benchmarks. The first rule is to select supervision by the reasoning operation it teaches, rather than by the actions, scenes or domains it shows. This runs against habit. Many teams assume that to improve a robotics task they should gather robotics-scene data, whereas WOVEN suggests that what drives transfer is which reasoning operation a sample exercises. The second rule is to prefer larger changes to the visual state, because they yield more robust capability. For the industry, the implication is that the unit of data curation should shift from domain to reasoning operation. It also implies that world-model training need not rely only on video generation: teaching a multimodal language model to reason about transitions directly is another route worth funding. One limit remains. Because the data comes from generative rollouts, the physical fidelity of those generators caps the quality of the supervision, and follow-up work will need to keep testing that ceiling.

Sources

FAQ

What is visual transition reasoning?

It is the ability to infer how a scene changes after an action is applied to it, for example where a pushed object ends up or whether a tower falls once a block is removed. WOVEN hypothesizes that failures in spatial, embodied, physical and temporal reasoning share this one deficit.

What are the key points of the WOVEN training recipe?

The paper reports two principles validated prospectively on held-out benchmarks. First, select supervision by the reasoning operation it teaches, not by the actions, scenes or domains it shows. Second, prefer examples with larger changes to the visual state, which gives more robust gains.

What limits should readers keep in mind?

The data comes from rollouts of video-pretrained generative models, so their physical fidelity bounds the supervision quality. The headline gains are self-reported across 26 external benchmarks, and 22 of them improved, not all. Teams should retest on their own tasks before relying on the numbers.