When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware figure
AlphaXiv 中文概览(可滚动查看)