Zhiyuan Wu, Zhuo Chen, Shan Luo
Vision tells a robot what an object is and roughly where. Touch tells it precisely what is happening at the single point where contact occurs. Combining them to answer "where on this object am I touching" is harder than it sounds, because the two views have to be aligned in space to a tolerance neither provides on its own.
The mismatch is fundamental rather than technical. A point cloud from a camera describes the object's surface in the camera's frame; a tactile reading describes pressure in the sensor's frame, with no inherent notion of where on the object that is. Fusing them means bridging that gap accurately enough that the answer is useful.
Getting it right matters for manipulation specifically. Knowing you have contact is not enough to decide whether a grip will hold, because whether it holds depends on where on the object you are gripping.
Vision and touch are complementary modalities essential for robotic perception and manipulation. While vision provides global object context, touch offers precise local information at contact points. Integrating these modalities for contact localization, i.e., predicting the location of touch on an object's surface, poses significant challenges due to the need for accurate spatial alignment between tactile data and visual geometry. To address this challenge, we propose…
Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints
arXiv (cs.AI) · July 17, 2026Candidate Attended Dialogue State Tracking Using BERT
arXiv (cs.LG) · July 17, 2026Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation
arXiv (cs.CV) · July 17, 2026PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.