EyeRobot 2.0: Active Gaze Lets Bimanual Robots Manipulate Precisely Without Wrist Cameras

Published · AI Daily — AI-assisted deep research, methodology & disclosure

EyeRobot 2.0 uses a single stereo camera and lets a robot swivel two eye viewpoints toward a 3D fixation point, as a human would. It gives more visual tokens to the image center. Hierarchical reinforcement learning controls the gaze, and gripper actions are expressed in a fixation-relative SE(3) frame. With passive stereo only, real-world success falls from 52% to 27%. EyeRobot 2.0 beats that baseline by 40% in real tests and 20% in simulation. It matches ego plus wrist cameras when wrist views are clear (69% vs. 64%) and more than doubles their success under occlusion (48% vs. 22%).

Background: why a robot needs to decide where to look

In fine-grained bimanual manipulation, the way a policy receives visual input sets a ceiling on what it can do. The common recipe pairs a head (ego) camera with a camera on each wrist. The wrist cameras sit close to the gripper and show grasp and insertion detail. They also carry costs: extra hardware, cables and calibration, and a harder problem. The object in the gripper often blocks the wrist view.

EyeRobot 2.0 takes a different path. Inspired by human vision, it uses a single stereo camera and lets the robot fixate on a 3D point in the scene, as a person does. The paper calls this mechanism Active Visual Fixation (AVF). The core bet is simple: instead of adding cameras to cover every angle, point one pair of cameras at the most important place at each moment.

Core architecture: three parts that work together

Physical fixation. The system picks a 3D fixation point and swivels the two eye viewpoints so both lines of sight converge on it. This is a real mechanical rotation, not a crop of the image.

Foveal processing. The images go into the network with more visual tokens allocated to the image centers and fewer to the periphery. This mirrors the human fovea, which has high resolution in the middle and low resolution at the edges. Compute is spent on task-relevant features.

Hierarchical gaze control. Gaze cannot move at random. It must stay in step with the task. The paper uses two levels. The low level is a gaze servoing policy, conditioned on a goal object, that aims the eyes at it. The high level is a target selector that emits the next fixation goal based on task progress.

Training: reinforcement learning on real-world data for both modules

The gaze servoing policy is trained with RL and a dense geometric reward. The reward measures how well the gaze aligns with the target. The signal is dense, which suits learning on a real robot.

The target selector is the more interesting piece. It is co-trained with the behavior cloning (BC) gripper policy. No human labels say where to look. The selector finds fixation sequences while it is optimized together with the gripper policy. The authors report that these sequences can resemble the fixation sequence of a human doing the same task. This suggests that gaze behavior can emerge from the goal of task success alone.

Action representation: a fixation-relative SE(3) frame

EyeRobot 2.0 also uses the fixation point to simplify action learning. It canonicalizes gripper information into a fixation-relative SE(3) frame. The gripper pose is no longer expressed relative to the robot base. It is expressed relative to the current fixation point.

This shrinks the action distribution the policy has to learn. The range of positions and orientations narrows, and the policy fits it more easily. For teleoperation datasets of limited size on real robots, this compression has real value.

Experiments and key results

The authors collected teleoperation data for 7 real-world tasks and 6 simulated tasks. They ran more than 1000 physical and 1800 simulated trials. The baselines were a passive stereo policy and an ego plus wrist camera policy, both trained on the same data. The main findings: - Removing wrist cameras is costly for standard policies. With only passive stereo, real-world success drops from 52% to 27%.

  • EyeRobot 2.0 closes this gap with only stereo. It beats passive stereo by 40% in real-world tests and by 20% in simulation.
  • When the wrist views are clear, it matches the ego plus wrist policies: 69% versus 64%.
  • When grasped objects occlude the wrist cameras, it reaches 48% while the ego plus wrist policies reach 22%. That is more than double.

The last result says the most. Wrist camera failure under occlusion is a physical limit. An active gaze camera sits outside the object and can keep watching the contact area from a useful angle.

Cost and latency tradeoffs

On the hardware side, the system removes the wrist cameras and keeps one stereo pair. It adds a mechanism that swivels the two viewpoints. The team trades camera count for mechanical degrees of freedom.

On the compute side, foveal token allocation concentrates compute on the image center, which in principle lowers the cost spent on the periphery. The abstract gives no numbers for latency or compute, so those gains must be checked against the full paper. Gaze servoing must also run continuously during the task, which adds control-loop complexity.

What this means for developers and industry

For robotics teams, this work offers a workable way to simplify sensors. Wrist camera cables, calibration, maintenance and occlusion are real costs in volume production and long-term deployment. If one swivelling stereo pair reaches equal or better success, the bill of materials gets shorter.

For learning methods, the paper shows the value of treating perception as action. Gaze is no longer a passive input. It is a decision that can be learned and optimized. The fixation-relative action representation can also be borrowed by other imitation learning frameworks.

Limits and future directions

Several points need care. First, the robot head needs a movable eye mechanism, which is not plug-and-play on current platforms. Second, 7 real-world tasks is a solid set but still lab scale, and generalization to open environments needs proof. Third, the target selector depends on joint training with the BC policy, so data quality will affect the quality of the fixation sequences. Fourth, how the system handles brief visual instability while the gaze shifts is a detail to check in the full paper.

Future directions include combining active gaze with large vision-language-action models, so that language instructions drive the fixation target. Others are testing on longer-horizon tasks and mobile manipulation, and designing lighter eye mechanisms. Overall, EyeRobot 2.0 makes a clear argument: teaching a robot where to look can replace adding more cameras.

Sources