Dual-scope surgical world modeling

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

Wanhao Liu, Jinsong Lin, Rulin Zhou, CHI KIT NG, Wenbin Pan, Zhiqing Tang, Dongyue Li, Miao Luo, Wu Yanshen, Panshuo Li, Zhiyong Xiong, Huxin Gao, Tamas Haidegger, Hongliang Ren*

Equal contribution. * Corresponding author.

CrossScope overview: Mother and Child scopes, directional evidence routing, and prediction gains.
Mother and Child scopes provide complementary wide- and near-field observations. CrossScope routes geometric motion from Mother to Child and pose-aligned appearance from Child to Mother.

Problem setting

Abstract

Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother-Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship.

We formulate Role-asymmetric dual-scope future prediction, where cross-view evidence is selectively transferred according to the prediction target and its spatial requirements. CrossScope is a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions.

Geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. On paired phantom and real-world ERCP episodes, CrossScope improves visual fidelity, structural preservation, target localization, and motion consistency over surgical video-generation baselines.

Role-asymmetric routing

Why role asymmetry?

The Mother scope supplies the stable global map, while the Child scope brings close-range detail from narrow bile ducts. Their fields of view, scale, and pose evolve independently during ERCP.

CrossScope turns this mismatch into two directed routes: M2C sends pose-derived motion evidence to the Child expert, and C2M injects Child appearance only where Mother-plane geometry supports the correspondence.

2 independently controlled scopes M2C geometry-guided motion C2M pose-aligned appearance
CrossScope qualitative prediction comparison across Mother and Child views.
Pose signal Pose signal used to align the dual-scope routing.
Motion route Motion experiment visualization for Mother-to-Child routing.
C2M support Spatial support diagnostic for Child-to-Mother appearance routing.

Model design

CrossScope architecture

Separate DiT experts read a shared pre-injection snapshot and receive only target-valid residual writes.

CrossScope dual-stream architecture with M2C geometric motion conditioning and C2M pose-aligned appearance conditioning.
Target-specific residual exchange. Pose Readout maps Mother states to a Mother-plane trajectory for M2C. C2M pose-aligns Current and causal History Child observations before writing or blending Mother residuals under geometric support.

Quantitative evidence

End-to-end results on the phantom benchmark

Frame fidelity across Child and Mother views. Higher PSNR/SSIM and lower FID/LPIPS are better.

ModelChild PSNRChild SSIMChild FIDChild LPIPSMother PSNRMother SSIMMother FIDMother LPIPS
HunyuanVideo-I2V34.3320.94218.2130.084835.1120.95334.3240.0352
Cosmos-H-Surgical35.1230.94317.4310.082135.7740.96933.6170.0347
Wan2.235.0030.94017.9860.085435.0010.95634.6180.0355
Early-Fusion Wan2.230.7900.91520.1900.087735.4320.96633.3780.0354
Symmetric Cross-Attention MoT33.3260.93118.3690.085236.2980.96629.0190.0351
CrossScope35.9060.95117.4020.081937.1470.97328.2320.0334

Child-scope motion quality

Papilla trajectory error on the phantom benchmark. Lower is better; values are in pixels.

ModelEndpoint error (px) ↓Centerline error (px) ↓
HunyuanVideo-I2V36.06819.126
Cosmos-H-Surgical28.88015.339
Wan2.235.09318.544
Early-Fusion Wan2.242.59021.733
Symmetric Cross-Attention MoT31.76712.414
CrossScope19.34311.238

CrossScope also achieves 0.948 Mother mask IoU, 0.839 Child papilla box IoU, and 0.927 detection recall.

Routing validity

Geometry-licensed C2M support

C2M has a different information contract from M2C: it writes only pose-aligned appearance from observed Child frames to the Mother expert. Current is the observed Child anchor; History contains strictly earlier cached observations.

Support is established before a residual is written. The predicted trajectory can select a Current-only, History-only, dual-source, or Null state without admitting future Child-token states as appearance evidence.

Four C2M spatial support cases showing the Mother anchor, Current coverage, causal History coverage, and final injection state.
Four C2M spatial-support cases: the routing state changes with the geometrically supported Current and causal History coverage.

Data coverage

Dual-scope observations across phantom and real-world settings

Representative Mother, mask, and Child observations make the coupled viewpoint relationship visible in both the phantom benchmark and real-world procedures.

Representative Mother, mask, and Child observations across phantom and real-world settings.
Paired dual-scope observations across phantom and real-world settings.

Reference

Citation

@misc{crossscope2026,
  title={CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction},
  author={Wanhao Liu and Jinsong Lin and Rulin Zhou and CHI KIT NG and Wenbin Pan and Zhiqing Tang and Dongyue Li and Miao Luo and Wu Yanshen and Panshuo Li and Zhiyong Xiong and Huxin Gao and Tamas Haidegger and Hongliang Ren},
  year={2026}
}