Zhiyuan Wu, Zhuo Chen, Shan Luo
Vision tells a robot what an object is and roughly where. Touch tells it precisely what is happening at the single point where contact occurs. Combining them to answer "where on this object am I touching" is harder than it sounds, because the two views have to be aligned in space to a tolerance neither provides on its own.
The mismatch is fundamental rather than technical. A point cloud from a camera describes the object's surface in the camera's frame; a tactile reading describes pressure in the sensor's frame, with no inherent notion of where on the object that is. Fusing them means bridging that gap accurately enough that the answer is useful.
Getting it right matters for manipulation specifically. Knowing you have contact is not enough to decide whether a grip will hold, because whether it holds depends on where on the object you are gripping.
Vision and touch are complementary modalities essential for robotic perception and manipulation. While vision provides global object context, touch offers precise local information at contact points. Integrating these modalities for contact localization, i.e., predicting the location of touch on an object's surface, poses significant challenges due to the need for accurate spatial alignment between tactile data and visual geometry. To address this challenge, we propose…
CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors
arXiv (cs.AI) · July 17, 2026Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
arXiv (cs.AI) · July 17, 2026A Formally Grounded ODRL Evaluator: Implementation and Comparison
arXiv (cs.LG) · July 17, 2026DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging