Selected Publications

# Equal contribution; * Corresponding author.

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction overview

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

Wanhao Liu#, Jinsong Lin#, Rulin Zhou#, Chi Kit Ng#, Wenbin Pan, Zhiqing Tang, Dongyue Li, Liwei Luo, Yanshen Wu, Panshuo Li, Zhiyong Xiong, Huxin Gao, Tamas Haidegger, Hongliang Ren*

arXiv preprint · 2026

Abstract

CrossScope studies role-asymmetric future prediction for Mother-Child endoscopic retrograde cholangiopancreatography, where two independently moving scopes provide complementary views without calibrated stereo geometry. Its dual-stream world model preserves view-specific experts and routes cross-view evidence according to target-specific spatial requirements, improving visual fidelity, structural preservation, target localization, and motion consistency.

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts overview

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

Rulin Zhou#, Wanhao Liu#, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren*

arXiv preprint · 2026

Abstract

Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument-tissue interactions. Surg-UniWorld introduces a Hierarchical Surgical Anchor, Anchor-Relative Modality Experts, and a Multimodal Control Expert to support coherent video generation under arbitrary combinations of edge, depth, and optical-flow controls. Experiments demonstrate improved generation quality, temporal consistency, and multimodal controllability.

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation overview

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation

Jinsong Lin#, Zikang Pan#, Wanhao Liu#, Chi Kit Ng#, Liangjing Shao, Zihang Yu, Ziyu Wang, Yin Wang, Jiaxi Wang, Jeremy Yuen-Chun Teoh, Zhiyong Xiong, Huxin Gao, Hongliang Ren*

arXiv preprint · 2026

Abstract

EndoWAM is a grounded World-Action Model for generalizable robotic endoscopic navigation. It predicts task-relevant target regions in future observations and couples a lightweight diffusion transformer with a discrete action expert through a shared predictive representation. The model enables real-time control and generalizes to unseen viewpoints, environments, and targets across multiple endoscopic procedures.