SurgCast predicts future instrument skeletons from the current skeleton, robot state, and prospective actions, then uses these structures to control surgical video generation. Dual-Path Semantic-Kinematic Injection combines geometric residuals with global-local kinematic modulation, while distribution-matching distillation transfers this interface to a four-step causal student. Experiments on SutureBot and SRTH-Porcine-Cholei evaluate visual quality and instrument control, with additional zero-shot transfer and physical master-device demonstrations.
CrossScope studies role-asymmetric future prediction for Mother-Child endoscopic retrograde cholangiopancreatography, where two independently moving scopes provide complementary views without calibrated stereo geometry. Its dual-stream world model preserves view-specific experts and routes cross-view evidence according to target-specific spatial requirements, improving visual fidelity, structural preservation, target localization, and motion consistency.
RoPE-Flow estimates robot 3D keypoints and joint angles from monocular images using graph-conditioned flow matching. A keypoint feature extractor combines positional cues with multi-scale appearance features, and graph-guided condition generation models correlations among keypoints. Bone-length and ray-depth losses promote topological consistency and scale stability. Evaluation covers different robot types and cameras, including a cross-scene benchmark for zero-shot robot pose estimation.
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument-tissue interactions. Surg-UniWorld introduces a Hierarchical Surgical Anchor, Anchor-Relative Modality Experts, and a Multimodal Control Expert to support coherent video generation under arbitrary combinations of edge, depth, and optical-flow controls. Experiments demonstrate improved generation quality, temporal consistency, and multimodal controllability.
EndoWAM is a grounded World-Action Model for generalizable robotic endoscopic navigation. It predicts task-relevant target regions in future observations and couples a lightweight diffusion transformer with a discrete action expert through a shared predictive representation. The model enables real-time control and generalizes to unseen viewpoints, environments, and targets across multiple endoscopic procedures.