SpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic Manipulation
SpatialHarness renders virtual views from a synced simulated scene for a frozen multimodal robot policy, with no fine-tuning or new sensors. Plug insertion rose from 26.7% to 66.7%, Tower of Hanoi from 0% to 100%.
Frontier multimodal foundation models such as GPT-6 Astra have recently shown real promise as direct robot controllers, yet their performance drops sharply once a task demands fine manipulation: inserting a plug, aligning two parts, or stacking objects with millimeter tolerance. The conventional explanation is that the policy is simply not capable enough, and the conventional remedies follow from it: gather more demonstrations, fine-tune the model, or switch to a stronger backbone. The SpatialHarness paper proposes a different diagnosis. A significant share of the failures, the authors argue, comes not from weak policy capability but from insufficient spatial observability. The spatial relationships that decide success, such as whether a plug is coaxial with its socket or whether a disk has cleared the top of a peg, may be poorly revealed by the existing physical camera setup, or hidden by the arm and the objects themselves. A model cannot reason about what it cannot see, and the same is true of a skilled human technician who is asked to work with one eye covered.
SpatialHarness answers this diagnosis with a test-time embodied harness. It does not fine-tune the policy and it does not change the physical sensing setup. Instead, it inserts a layer of spatial scaffolding between the frozen multimodal policy and the real world. The system maintains an online simulated scene that is synchronized with real-world execution, identifies the spatial relationships that are critical for the current task, and renders complementary virtual views that expose those relationships to the policy. The real cameras supply the facts, and the virtual cameras re-present those facts from angles the model can read more easily. The design has a clear advantage: a simulated camera can sit where no physical camera could, for example looking along the inside of a socket or across the thin gap between two parts, and it costs nothing in hardware. For closed models that can only be reached through an API, an approach that improves the input and leaves the weights untouched is especially practical.
Keeping the simulated scene aligned with reality is the hardest part of the method. Once an object is grasped, moved and released, pose-estimation error can accumulate, and a virtual view that disagrees with the real scene may mislead the policy more than no extra view at all. The authors therefore develop interaction-aware scene synchronization, which distinguishes three modes: static, held, and transition. Based on the abstract, the natural reading is that a static object can be trusted to stay where perception placed it, a held object moves with the gripper and its relation to the end effector is roughly fixed, and a transition, such as the moment of grasping or releasing, is the point at which the state is changing and the scene must be realigned with more care. The exact estimation and correction algorithms used in each mode should be checked against the full paper. The principle, however, deserves attention: treat the current stage of interaction as a prior that tells the system which source of information to trust, rather than applying one tracking rule to every moment of a task.
The evaluation covers four real-robot manipulation tasks that span precise geometric alignment, object-relative placement, and articulated-object interaction. With the same frozen GPT-6 Astra policy, SpatialHarness substantially improves task success. Plug insertion rises from 26.7% to 66.7%, and Tower of Hanoi rises from 0% to 100%. The two numbers carry different meanings. A jump from zero to perfect on Tower of Hanoi suggests that the bottleneck was almost entirely about seeing the relative positions of disks and pegs; once those relations were shown clearly, the model's existing reasoning and planning were enough. Plug insertion is closer to the difficulty of real industrial work. Success more than doubles, yet roughly one attempt in three still fails, which indicates that observability is one bottleneck among several, with contact handling and force control still in play. It is also worth keeping a sober view of the evidence: there are only four tasks, and the number of trials and the breakdown of failure modes should be checked in the paper and on the project website.
The broader significance lies in how the work re-attributes capability. For several years the dominant story in robot learning has been to train larger models on more data. SpatialHarness suggests that strong foundation models may already hold a good part of the skill needed for fine manipulation, and that poor observation conditions are keeping it locked away. Test-time augmentation needs no retraining, can be layered on top of any frozen backbone, and can be reused when the backbone is upgraded. It also has clear limits. The method depends on the quality of the simulated scene, which needs object models and reliable pose estimation. The abstract says nothing about deformable or unseen objects, and it does not report the latency or compute cost of online rendering. For the industry, the paper brings a layer of engineering into view: the harness built around a frozen model. As foundation models converge in raw capability, real-robot performance may depend more and more on who gives the model the better way to look at the world.
Sources
FAQ
What core problem does SpatialHarness address?
It argues that frontier multimodal models often fail at fine manipulation not because the policy is weak, but because task-critical spatial relationships are poorly visible in the existing camera views. SpatialHarness adds virtual views at test time, with no fine-tuning and no change to the physical sensors.
Why does the scene synchronization distinguish static, held, and transition modes?
When an object is grasped, moved or released, a simulated scene that drifts from reality would mislead the policy. Separating the three modes lets the system decide which source of information to trust at each stage of the interaction. The exact algorithms per mode are in the full paper.
What do the results show, and what are the limits?
With the same frozen GPT-6 Astra policy on four real-robot tasks, plug insertion rose from 26.7% to 66.7% and Tower of Hanoi from 0% to 100%. That suggests better spatial observability unlocks existing skill. The limits are the small task count, the reliance on simulation quality, and undisclosed latency and compute cost.