OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation

Conference on Neural Information Processing Systems (NeurIPS) 2025

Raktim Gautam Goswami1, Prashanth Krishnamurthy1, Yann LeCun2,3, Farshad Khorrami1

1 New York University Tandon School of Engineering
2 New York University Courant Institute of Mathematical Sciences
3 Meta-FAIR
Fig. 1: OSVI-WM infers the task from the expert demonstration and, along with the agent’s observation “foresees” future latent states using a world-model-guided trajectory generation module. The predicted trajectory is decoded into physical waypoints for control.

OSVI-WM Contributions

Video

Results

Table 1: Success rates (in %) comparison on the Meta-World and Pick-and-Place simulation benchmarks. Best results are highlighted. We also report if a method uses additional training data.
Table 2: Real-World experiments: Success rates and execution breakdowns (in %) are reported. T-OSVI* denotes T-OSVI [10] aided with end-effector depth sensing for improved grasping.

Experimental result clips for all the benchmarks are available in the project video.

Citation

@article{goswami2025osvi,
  title={Osvi-wm: One-shot visual imitation for unseen tasks using world-model-guided trajectory generation},
  author={Goswami, Raktim Gautam and Krishnamurthy, Prashanth and LeCun, Yann and Khorrami, Farshad},
  journal={arXiv preprint arXiv:2505.20425},
  year={2025}
}