In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use figure
AlphaXiv 中文概览(可滚动查看)