Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother-Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship.
We formulate Role-asymmetric dual-scope future prediction, where cross-view evidence is selectively transferred according to the prediction target and its spatial requirements. CrossScope is a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions.
Geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. On paired phantom and real-world ERCP episodes, CrossScope improves visual fidelity, structural preservation, target localization, and motion consistency over surgical video-generation baselines.