Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
Authors: Meng Chu, Zhedong Zheng,
Wei Ji,
Tingyu Wang,
Tat-Seng Chua
Published in European conference on computer vision (ECCV), 2024
Recommended citation: Meng Chu, Zhedong Zheng, Wei Ji, Tingyu Wang, Tat-Seng Chua, "Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching." European conference on computer vision (ECCV), 2024.
Download PDF: https://zdzheng.xyz/files/2024/ECCV24-GeoText.pdf
中文解读: https://www.zhihu.com/question/660698707/answer/3575966275
Code is available at: https://multimodalgeo.github.io/GeoText/
Abstract: Navigating drones through natural language commands remains challenging due to the dearth of accessible multi-modal datasets and the stringent precision requirements for aligning visual and textual data. To address this pressing need, we introduce GeoText-1652, a new natural language-guided geo-localization benchmark. This dataset is systematically constructed through an interactive human-computer process leveraging Large Language Model (LLM) driven annotation techniques in conjunction with pre-trained vision models. GeoText-1652 extends the established University-1652 image dataset with spatial-aware text annotations, thereby establishing one-to-one correspondences between image, text, and bounding box elements. We further introduce a new optimization objective to leverage fine-grained spatial associations, called blending spatial matching, for region-level spatial relation matching. Extensive experiments reveal that our approach maintains a competitive recall rate comparing other prevailing cross-modality methods. This underscores the promising potential of our approach in elevating drone control and navigation through the seamless integration of natural language commands in real-world scenarios.
@inproceedings{GeoText1652,
author = "Chu, Meng and Zheng, Zhedong and Ji, Wei and Wang, Tingyu and Chua, Tat-Seng",
title = "Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching",
abstract = "Navigating drones through natural language commands remains challenging due to the dearth of accessible multi-modal datasets and the stringent precision requirements for aligning visual and textual data. To address this pressing need, we introduce GeoText-1652, a new natural language-guided geo-localization benchmark. This dataset is systematically constructed through an interactive human-computer process leveraging Large Language Model (LLM) driven annotation techniques in conjunction with pre-trained vision models. GeoText-1652 extends the established University-1652 image dataset with spatial-aware text annotations, thereby establishing one-to-one correspondences between image, text, and bounding box elements. We further introduce a new optimization objective to leverage fine-grained spatial associations, called blending spatial matching, for region-level spatial relation matching. Extensive experiments reveal that our approach maintains a competitive recall rate comparing other prevailing cross-modality methods. This underscores the promising potential of our approach in elevating drone control and navigation through the seamless integration of natural language commands in real-world scenarios.",
booktitle = "European conference on computer vision (ECCV)",
code = "https://multimodalgeo.github.io/GeoText/",
url = "https://zdzheng.xyz/files/2024/ECCV24-GeoText.pdf",
blog = "https://www.zhihu.com/question/660698707/answer/3575966275",
funding = "SRG2024-00002-FST",
year = "2024" }