TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging figure
AlphaXiv 中文概览(可滚动查看)