S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving figure
AlphaXiv 中文概览(可滚动查看)