Keypoint tracking alone is insufficient for object interaction tasks such as sitting on a chair, wiping a board, or pushing furniture, where the robot can reach the correct pose without making meaningful physical contact with the object. We present ContactMimic, a learning framework that tracks explicit part-level binary contact commands alongside keypoint trajectories. ContactMimic is made possible through the use of contact-following rewards and a trajectory augmentation scheme aimed at breaking the correlations between keypoint trajectories and contact labels. The resulting policy successfully decouples contact behavior from keypoint geometry, and achieves precise physical contact as well as contact-controllability (produce or suppress contact during deployment as desired). Simulation experiments across 10 diverse human-object interaction motions confirm that ContactMimic exhibits contact controllability that enables it to complete manipulation tasks without task-specific rewards, while also outperforming keypoint-only trackers on contact-relevant tasks. Ablations confirm the necessity of the proposed trajectory augmentation scheme and sim2real deployment validates contact controllability in the real world across 5 different motions.
@article{li2026contactmimic,title={ContactMimic: Humanoid Object Interaction via Contact Control},author={Li, Xinyao and He, Xialin and Dong, Runpei and Gupta, Saurabh},journal={arXiv 2026},year={2026},month=jul,note={Xinyao Li and Xialin He contributed equally. University of Illinois Urbana-Champaign.}}
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
ULTRA is an all-in-one controller for humanoid loco-manipulation: track when references exist; act from egocentric perception and sparse intent when they don’t.
@inproceedings{he2026ultra,title={ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation},author={He, Xialin and Xu, Sirui and Li, Xinyao and Dong, Runpei and Bian, Liuyu and Wang, Yu-Xiong and Gui, Liang-Yan},booktitle={IROS 2026},year={2026},month=mar,note={Xialin He and Sirui Xu contributed equally.}}
2025
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
Propose Camera Depth Models (CDMs) as a simple plugin on daily-use depth cameras, which take RGB images and raw depth signals as input and output denoised, accurate metric depth.
@inproceedings{liu2026manipulation,title={Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots},author={Liu, Minghuan and Zhu, Zhengbang and Han, Xiaoshen and Hu, Peng and Lin, Haotong and Li, Xinyao and Chen, Jingxiao and Xu, Jiafeng and Yang, Yichu and Lin, Yunfeng and Li, Xinghang and Yu, Yong and Zhang, Weinan and Kong, Tao and Kang, Bingyi},booktitle={ICLR 2026},year={2025},month=sep,note={Minghuan Liu, Zhengbang Zhu, Xiaoshen Han and Peng Hu contributed equally.}}
RHINO: Learning Real-Time Humanoid-Human-Object Interaction from Human Demonstrations
Propose the first real-time humanoid interaction framework capable of learning from human demonstrations, enabling dynamic task-switching and immediate responses to human instructions.
@inproceedings{chen2026rhino,title={RHINO: Learning Real-Time Humanoid-Human-Object Interaction from Human Demonstrations},author={Chen, Jingxiao and Li, Xinyao and Cao, Jiahang and Zhu, Zhengbang and Dong, Wentao and Liu, Minghuan and Wen, Ying and Yu, Yong and Zhang, Liqing and Zhang, Weinan},booktitle={IROS 2026},year={2025},month=feb,note={Jingxiao Chen, Xinyao Li and Jiahang Cao contributed equally.}}