STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision figure
AlphaXiv 中文概览(可滚动查看)