A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data figure
AlphaXiv 中文概览(可滚动查看)