
SurgCast: Action-Conditioned Future Skeletons for Controllable Surgical World Models
arXiv preprint · 2026
摘要
SurgCast predicts future instrument skeletons from the current skeleton, robot state, and prospective actions, then uses these structures to control surgical video generation. Dual-Path Semantic-Kinematic Injection combines geometric residuals with global-local kinematic modulation, while distribution-matching distillation transfers this interface to a four-step causal student. Experiments on SutureBot and SRTH-Porcine-Cholei evaluate visual quality and instrument control, with additional zero-shot transfer and physical master-device demonstrations.




