LEGO: Lifting-Free Exocentric-to-Egocentric Video Generation, Where a Synthesizer Supplies Structure and Diffusion Supplies Detail

Published · AI Daily — AI-assisted deep research, methodology & disclosure

LEGO (arXiv 2610.12442) turns one exocentric video into an egocentric one without depth or point clouds. An LVSM-style transformer renders the structure, a diffusion model restores detail, and correspondence confidence guides denoising.

Generating an egocentric video from a single exocentric recording is one of the hardest cases of novel view synthesis. The reason is simple. The two cameras share very little overlap, so most of the target view never appears in the input. The payoff is large, though. Robot imitation learning needs demonstrations seen from the actor's own viewpoint. AR, VR and skill coaching need first-person footage. Collecting that footage at scale is slow and expensive, while third-person video is available almost everywhere. On 8 October 2026, Suhwan Cho and four co-authors posted LEGO to arXiv (2610.12442, cs.CV), and it proposes a route that departs from the current mainstream. Start with what the strongest existing methods do. According to the abstract, the state of the art reconstructs the scene explicitly. It estimates depth, lifts the video into a point cloud, re-renders that cloud from the egocentric camera, and passes the result to a video diffusion model as a condition. This is a deterministic mapping, because every pixel lands on exactly one reprojected location. The upside is that texture survives well. The downside is just as clear: any error in the depth estimate becomes misplaced content in the condition. When two cameras barely overlap, depth errors are almost unavoidable, so the quality of the whole pipeline is capped by the quality of its geometry.

LEGO asks a different question: what should a video diffusion model actually receive as its condition? Its answer is lifting-free. The authors fine-tune an LVSM-style transformer, which they call a learned view synthesizer, to render the egocentric view directly. It uses no depth, no point cloud and no reprojection. The network resolves cross-view correspondence internally. Where the explicit pipeline is deterministic, this mapping is probabilistic. For each region, the synthesizer averages over candidate source locations, weighted by a learned correspondence distribution. The result keeps structure and loses fine texture, because averaging smooths away detail. The rendered view looks softer than a reprojection, yet its content sits roughly where it should.

The authors do not hide this cost. They turn it into the design rationale. Their argument is that the trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. In other words, the condition has to say where things are and what the overall scene looks like, while the generator fills in hair, grain and surface texture on its own. The final sentence of the abstract states the division of labor: the synthesizer supplies view structure, and the diffusion model supplies detail. Many coarse-to-fine systems follow a similar logic. What is new here is that the coarse stage is no longer a hand-built geometric pipeline. It is a learned module trained end to end for the task.

LEGO also extracts more value from the distribution itself. A concentrated correspondence distribution means the synthesizer is confident about where a region comes from. A diffuse one means many candidate locations compete and the result is less reliable. The paper converts this concentration into a per-region confidence and uses it twice. First, it masks low-confidence regions, so unreliable condition signals are not forced on the generator. Second, during the early denoising steps, which form the global layout, it guides the generator toward high-confidence areas. The overall composition settles first on solid ground, and the remaining regions follow. The logic is natural: listen to the condition more where it is sure, and less where it is not.

On results, the abstract states that LEGO consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The abstract gives no metric values, so the size of the gain can only be judged from the experiments section of the full paper. Two cautions are worth keeping in mind. First, regions that were occluded or never observed must still be imagined by the generator. A confidence score makes the system more honest about where it can be trusted, but it cannot create information that the input never contained. Second, how far the cross-dataset claim reaches depends on which datasets the authors tested, and that needs checking against the paper itself.

Even with those caveats, the value of LEGO reaches beyond this single task. It states a design principle that other work can reuse: tailor the condition to the strengths of the generator instead of chasing a sharp, physically exact input. Tasks that drive diffusion models with geometric conditions, such as video editing, monocular re-rendering and robot world models, face the same tension between exact but brittle geometry and soft but robust structure. For practitioners, the lesson is concrete. If your pipeline fails because an upstream estimator makes confident mistakes, consider replacing the hard mapping with a probabilistic one, keep its uncertainty, and let the diffusion model handle the detail. LEGO shows that this swap can work for one of the least constrained view-synthesis problems in video.

Sources

FAQ

What does "lifting-free" mean in LEGO?

LEGO does not estimate depth, build a point cloud or reproject pixels. It fine-tunes an LVSM-style transformer to render the egocentric view directly, and the network resolves cross-view correspondence internally. That rendering then serves as the condition for a video diffusion model.

Why can a blurrier condition help a diffusion model?

The authors argue that denoising training already excels at restoring detail, so the condition should prioritize structural alignment over sharpness. The probabilistic mapping averages each region over candidate source locations, which keeps structure and loses fine texture. The explicit pipeline keeps texture but turns depth errors into misplaced content.

How is the per-region confidence computed and used?

The synthesizer learns a correspondence distribution for each region. A concentrated distribution means the model is sure where the content comes from. The paper uses this concentration as a confidence score, first to mask low-confidence regions and second to guide the generator toward high-confidence areas during the early denoising steps that form the layout.