JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling figure
AlphaXiv 中文概览(可滚动查看)