Convergent binocular stereo: Depth perception for humanoid robot vision.
Journal:
Science robotics
Published Date:
Aug 26, 2026
Abstract
The design of robotic binocular camera systems has been inspired by human vision, as has their use in humanoid robots. One aspect of this inspiration has yet to play a major role, namely, that humans determine depth using a convergent binocular imaging geometry with both eyes pointing at the same location. Although some robot heads have the functionality to use vergence and version movements and thus alter binocular imaging geometry, a method to exploit it for depth computation has not been explored. Even though to an observer, their motions and appearance might resemble that of humans, in detail, the resemblance to the motions required for actual convergent depth computation is not seen. To bridge this gap, we present convergent binocular stereo (CBS), a stereo algorithm designed to provide a foundation for purposeful binocular computations under the humanoid constraint, intended for active binocular robots. CBS computes horizontal and vertical disparities using a coarse-to-fine refinement strategy with Gabor-filtered responses. We also introduce the Convergent Binocular Stereo-BenchMark (CBS-BM), a convergent, natural image dataset containing 49 scenes with ground-truth horizontal disparity collected on a four-degrees-of-freedom robotic system. Our evaluation, a quantitative comparison between parallel and convergent stereo systems, shows that CBS is broadly competitive with state-of-the-art parallel methods, even outperforming them in scenes with repeated patterns and in mean horizontal disparity and depth error over all scenes. Although not intended to replace parallel stereo where humanlike behavior is unnecessary, CBS enables functionally realistic depth computation for humanoid robotic heads.
Authors
Keywords
No keywords available for this article.