WorldAlign: A Decoupled 4D Reward for World-Consistent Video Generation

Published · AI Daily — AI-assisted deep research, methodology & disclosure

WorldAlign splits the reward in two. Static regions align to a geometric prior by masked reprojection, plus a camera-motion reward. A VLM judge scores dynamic subjects with checklists. No human labels. Wan2.1 and Wan2.2 improve.

Video generation has advanced mostly along one axis: looking real. Textures are finer, lighting is smoother, and camera paths are steadier. But the question changes once people treat these models as world simulators, for robot rehearsal, game scenes, or training environments for embodied agents. A simulator must do more than look good. It must keep its own rules. A wall seen from one side must keep its shape when the camera circles back. A person who walks for ten seconds must not quietly change face or clothes. The WorldAlign paper, by Jing He, Kaixin Ding, Xingye Tian, Guibao Shen, Wenhang Ge and colleagues, names this requirement 4D world consistency. The three spatial dimensions must stay coherent, and so must motion and appearance along the time axis. The first part is static consistency. The second is dynamic consistency. The paper was published on 8 October 2026 under cs.CV.

A natural way to improve consistency is geometry-aware post-training. The generator makes a video. A geometry model then estimates depth and camera pose, reprojects one frame into another, and measures the error. That error becomes a reward for tuning the generator. This route has two known weaknesses. First, it usually assumes the scene is static. When a person walks or a car drives, reprojection reads real motion as geometric error, and the reward punishes exactly the dynamics a good video needs. Second, even methods that tolerate dynamic scenes give unreliable static-consistency feedback, because moving and static pixels are mixed into one error. Dynamic consistency, meaning plausible motion and stable appearance over time, is often ignored or measured with crude proxies.

The central idea of WorldAlign is simple. Static and dynamic content follow different assumptions, so they should not be measured with the same ruler. The framework first separates the frame semantically into static regions and dynamic subjects. Each part is then aligned with the world prior that fits its assumptions. For static regions the prior is geometric. A semantically guided, masked reprojection computes cross-view geometric alignment only on pixels judged static. Moving subjects are excluded, so the feedback is cleaner and more reliable. The authors add one practical patch. A pure geometry reward has a shortcut: if the video barely moves, the reprojection error is small. The optimizer finds this. So an auxiliary camera-motion reward penalizes near-static solutions and forces real viewpoint change while geometry stays consistent. This is a typical defense against reward hacking. It is also a reminder that any single computable metric will be gamed by a strong optimizer.

For dynamic subjects, geometry cannot help. Whether motion is plausible is a question of common sense and physics, not of geometry. The authors use a strong vision-language model as the dynamic world prior and let it act as a judge. The key design is the sample-specific checklist. For each generated video, the system builds concrete questions and scores them item by item across four aspects: dynamicity, physical plausibility, shape consistency, and texture consistency. Dynamicity asks whether things that should move do move. Physical plausibility asks whether actions respect gravity, contact and inertia. Shape and texture checks watch whether a subject deforms or changes skin over time. A checklist is better than one vague score. It anchors the judge to specific evidence and makes the reward easier to inspect and debug. Just as important, none of this needs human preference annotations. Both rewards come from priors, so online post-training is possible and data cost drops.

The paper tests the method on two pretrained image-to-video generators, Wan2.1 and Wan2.2. It reports that WorldAlign improves static and dynamic consistency together, and that it does so without suppressing overall motion or subject motion. That last point matters. Many consistency gains in the literature come from making videos duller, and a method that wins by freezing the scene is not a win. From the abstract alone we cannot see exact numbers, ablation details or compute cost. Readers should check the full paper and the project page at worldalign.github.io for those. Several open questions remain from the method design. A VLM judge has its own bias and instability, and the reward can be no better than the judge. If semantic segmentation mislabels static and dynamic regions, masked reprojection will inject noise. Geometric priors can also fail on reflective or transparent surfaces and at extreme viewpoints.

Even so, the principle is valuable. When a world is made of parts with different natures, the reward should be split along the same lines, and each part should be aligned with a prior suited to it. This idea may extend to longer videos, multi-subject interaction, and eventually to visual world simulation that engineers can rely on. For teams building video models, the practical lesson is to audit your reward for hidden assumptions, such as a static scene, and to guard each reward term against the cheapest way to satisfy it.

Sources

FAQ

What does decoupled mean in WorldAlign?

The frame is split by semantics into static regions and dynamic subjects, and each gets its own world prior. Static regions are checked against a geometric prior for cross-view 3D structure. Dynamic subjects go to a vision-language model that checks motion, physics, shape and texture. The two rewards do not interfere.

Why add a camera-motion reward?

A static-consistency reward has a shortcut. If the video barely moves, the reprojection error is small. A model can learn to produce near-static clips. The auxiliary camera-motion reward penalizes that solution and pushes the model to change viewpoint while staying geometrically consistent.

Why does it need no human preference annotations?

The static reward comes from a geometric prior through computable reprojection alignment. The dynamic reward comes from a strong VLM that scores sample-specific checklists. Neither needs human good-versus-bad pairs, so online post-training is possible. The cost is that reward quality depends on how reliable these priors are.