LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models figure
AlphaXiv 中文概览(可滚动查看)