GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views figure
AlphaXiv 中文概览(可滚动查看)