4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench asks an agent to watch a video and write an executable graphics program that reproduces the scene structure and its motion. The scenes cover deformation, fluid flow and fracture, using both real videos and synthetic scenes. The authors benchmarked frontier models and found that strong static reconstruction does not yet mean reliable reconstruction of complex dynamics. The benchmark turns understanding how the world moves into a task that can be run and scored, and it gives the field a public testbed for tracking progress.
4DCodeBench is a new benchmark for 4D inverse graphics through code generation. Here 4D means three spatial dimensions plus time. The task is simple to state. An agent receives a video and must write an executable graphics program. When the program runs, it must reproduce the scene structure and the motion seen in the video. This differs from most video-understanding benchmarks. The answer is not a caption or a class label. It is code. Code can be run, inspected and compared with the source video frame by frame, so the evaluation is harder to bluff. Some background helps. Forward graphics turns a scene description into an image: given geometry, materials, lighting and physical rules, a renderer produces pixels. Inverse graphics runs the other way: given pixels, recover the scene description behind them. Earlier work leaned on differentiable rendering, neural radiance fields or Gaussian splatting. Those methods output large sets of parameters that are hard to read or edit. 4DCodeBench takes a different route. It asks for a compact, programmatic representation. The agent must turn visual observations into abstractions of scene structure and dynamics. Where needed, it must implement abstractions such as physical simulation to reproduce complex behavior. In short, the test is not what the agent sees. The test is whether it understands why the scene moves as it does.
On data, the paper's abstract says the authors curated a set of real-world videos and built synthetic scenes that span diverse physical phenomena, including deformation, fluid flow and fracture. The two sources serve different goals. Real videos test generalization under noise, occlusion and cluttered backgrounds. Synthetic scenes give controlled ground truth. The designers know the underlying physical parameters and the generating program, so they can locate exactly where a model fails. The three named phenomena are also a deliberate spread. Deformation points to elasticity, fluid flow to continuum mechanics, and fracture to discontinuous failure. Their numerical methods, parameter meanings and visual signatures differ a lot. One recipe is unlikely to cover all three.
The abstract does not describe the evaluation protocol, so the next paragraph is a general inference about tasks of this type. Check the full paper for the real details. A typical loop looks like this: perceive the video, propose a scene and a physical hypothesis, write graphics code, run it, compare the render with the target, then revise the code from the differences. The comparison can look at per-frame appearance and at temporal consistency across frames. Appearance mostly reflects static reconstruction. Temporal consistency is where dynamics show up. Every step can fail. The agent may pick the wrong physical model, set parameters at the wrong scale, choose a time step that makes the simulation diverge, or write code that does not run at all. The central finding is stated in the abstract. The authors benchmarked frontier models extensively and found that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. This is informative. It suggests that a model can place objects, materials and cameras correctly, and still fail to choose the right physical mechanism, estimate its parameters, and integrate it correctly over time. A static scene is like a well-drawn set. A dynamic scene demands a grasp of cause and effect. The abstract gives no scores, and this article does not invent any. For model rankings and per-phenomenon gaps, read the paper and the public benchmark page. Cost and latency are also absent from the abstract, so what follows is reasoned speculation. Code generation, simulation runs and several revision rounds add up, and the total grows with the number of iterations. Fluid and fracture simulations carry real compute cost, and every trial by the agent means executing a program. When teams compare systems on this benchmark, they should report iteration counts, token use and simulation time next to the final score. Otherwise a stronger result may only mean a larger budget. For developers and enterprises, three points stand out. First, a programmatic representation is editable, reusable and can plug into a physics engine. That matters for digital twins, robot simulation, game and film effects, and scientific visualization. A video that turns into a simulation script you can tune and rerun is far more useful than a block of opaque weights. Second, the benchmark offers another way to test world models. Whether a model understands physics can be checked by asking whether it can write a simulation that runs and matches. Third, the benchmark is public. The community can track progress with one yardstick, and teams can tell whether perception, planning, coding or physical knowledge is the bottleneck.
The limits are real. The abstract carries little detail, so scoring rules, the model list and the task count need checking in the full text. Visual similarity is not physical correctness: different programs can produce videos that look alike, and a metric must separate a lucky match from real understanding. Synthetic scenes differ from real footage, so gains on one may not carry over to the other. Results also depend on the graphics libraries and simulators the agents use, and familiarity with a given library can leak into the score. Looking ahead, several directions deserve attention: agents that use simulator feedback to correct themselves, training on data built for physical program synthesis, combining differentiable physics with code generation, and stricter metrics for parameter estimation. The value of 4DCodeBench is that it turns understanding a dynamic world into an engineering problem that can be run and scored. Today's frontier models still have a clear distance to cover on that road.