Homanga Bharadhwaj, Yash Jangir
Watch someone reach for a cup and you already know roughly what happens next. It will be lifted, not slid. A drawer will come out toward you. A lid will rotate shut. You are not calculating physics, you are drawing on a lifetime of having seen objects behave.
That anticipation is what a robot needs before it can act sensibly, and this work learns it from ordinary single-camera videos of people handling things. Given a short observed clip, MotionForesight predicts the future 3D trajectories of points on the object being manipulated.
Framing it as object-centred 3D prediction rather than pixel prediction is the useful move. What matters for acting is where the thing will be in space, not what the next frame will look like, and monocular video of humans doing everyday tasks is an abundant source of exactly that lesson.
Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act in the real world. We study how to learn this anticipation from ordinary monocular videos of human-object interaction. Given a short observed video context, MotionForesight predicts future 3D trajectories for points on the manipulated object. This casts…
A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
arXiv (cs.CV) · July 17, 2026Vision-Language Assistant for Emotional Reactions to Risky Driving
arXiv (cs.LG) · July 17, 2026Cluster-Aware Matching via Laplacian Optimal Transport
arXiv (cs.LG) · July 17, 2026Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems