Fetching the paper…
Reading the bibliography…
Navigating drones through natural language commands remains challenging due to the dearth of accessible multi-modal datasets and the stringent precision requirements for aligning visual and textual data.
Meguro, J.I., Ishikawa, K., Hasizume, T., Takiguchi, J.I., Noda, I., Hatayama, M.: Disaster information collection into geographic information system using rescue robots. In: IROS. pp. 3514–3520 (2006)
2006
Earlier work this paper cites.
Chen, X., Fang, H., Lin, T.Y., Vedantam, R., Gupta, S., Dollar, P., Zitnick, C.L.: Microsoft coco captions: Data collection and evaluation server. arXiv (2015)
2015
Earlier work this paper cites.
Doersch, C., Gupta, A., Efros, A.A.: Unsupervised visual representation learning by context prediction. In: ICCV. pp. 1422–1430 (2015)
2015
Earlier work this paper cites.
Workman, S., Souvenir, R., Jacobs, N.: Wide-area image geolocalization with aerial reference imagery. In: ICCV. pp. 1–9 (2015)
2015
Earlier work this paper cites.
Brunsting, S., De Sterck, H., Dolman, R., van Sprundel, T.: Geotexttagger: High-precision location tagging of textual documents using a natural language processing approach. arXiv (2016)
2016
Earlier work this paper cites.
Chandarana, M., Meszaros, E.L., Trujillo, A., Allen, B.D.: ’fly like this’: Natural language interface for uav mission planning. In: ACHI (2017)
2017
Earlier work this paper cites.
Jin Kim, H., Dunn, E., Frahm, J.M.: Learned contextual feature reweighting for image geo-localization. In: CVPR. pp. 2136–2145 (2017)
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: ICLR (2017)
2017
Earlier work this paper cites.
Anderson, P., et al. : Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: CVPR (2018)
2018
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL (2019)
2019
Earlier work this paper cites.
Huang, B., Bayazit, D., Ullman, D., Gopalan, N., Tellex, S.: Flight, camera, action! using natural language and mixed reality to control a drone. In: ICRA. pp. 6949–6956 (2019)
2019
Earlier work this paper cites.
Li, K., Zhang, Y., Li, K., Li, Y., Fu, Y.: Visual semantic reasoning for image-text matching. In: ICCV. pp. 4654–4662 (2019)
2019
Earlier work this paper cites.
Liu, L., Li, H.: Lending orientation to neural networks for cross-view geo-localization. In: CVPR (2019)
2019
Earlier work this paper cites.
Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., Savarese, S.: Generalized intersection over union: A metric and a loss for bounding box regression. In: CVPR. pp. 658–666 (2019)
2019
Earlier work this paper cites.
Shi, Y., Liu, L., Yu, X., Li, H.: Spatial-aware feature aggregation for image based cross-view geo-localization. In: NeurIPS. vol. 32 (2019)
2019
Earlier work this paper cites.
Wang, Z., Liu, X., Li, H., Sheng, L., Yan, J., Wang, X., Shao, J.: Camp: Cross-modal adaptive message passing for text-image retrieval. In: ICCV. pp. 5764–5773 (2019)
2019
Earlier work this paper cites.
Yu, Q., Wang, C., Cetiner, B., Yu, S.X., Mckenna, F., Taciroglu, E., Law, K.H.: Building information modeling and classification by visual learning at a city scale. NeurIPS 30
2019
Earlier work this paper cites.
Blukis, V., Terme, Y., Niklasson, E., Knepper, R.A., Artzi, Y.: Learning to map natural language instructions to physical quadcopter control using simulated flight. In: CoRL. pp. 1415–1438 (2020)
2020
Earlier work this paper cites.
Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., Liu, J.: Uniter: Universal image-text representation learning. In: ECCV. pp. 104–120 (2020)
2020
Earlier work this paper cites.
Hao, W., Li, C., Li, X., Carin, L., Gao, J.: Towards learning a generic agent for vision-and-language navigation via pre-training. In: CVPR (2020)
2020
Earlier work this paper cites.
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al.: Oscar: Object-semantics aligned pre-training for vision-language tasks. In: ECCV. pp. 121–137 (2020)
2020
Earlier work this paper cites.
Majumdar, A., Shrivastava, A., Lee, S., Anderson, P., Parikh, D., Batra, D.: Improving vision-and-language navigation with image-text pairs from the web. In: ECCV (2020)
2020
Earlier work this paper cites.
Rashid, M.T., Zhang, D.Y., Wang, D.: Socialdrone: An integrated social media and drone sensing system for reliable disaster response. In: INFOCOM. pp. 218–227 (2020)
2020
Earlier work this paper cites.
Thomason, J., Gordon, D., Bisk, Y.: Vision-and-dialog navigation. In: CoRL (2020)
2020
Earlier work this paper cites.
Vaucher, A.C., Zipoli, F., Geluykens, J., Nair, V.H., Schwaller, P., Laino, T.: Automated extraction of chemical synthesis actions from experimental procedures. Nature communications 11
2020
Earlier work this paper cites.
Zhang, Q., Lei, Z., Zhang, Z., Li, S.Z.: Context-aware attention network for image-text retrieval. In: CVPR. pp. 3536–3545 (2020)
2020
Earlier work this paper cites.
Zheng, Z., Wei, Y., Yang, Y.: University-1652: A multi-view multi-source benchmark for drone-based geo-localization. In: ACM MM. pp. 1395–1403 (2020)
2020
Earlier work this paper cites.
Zheng, Z., Zheng, L., Garrett, M., Yang, Y., Xu, M., Shen, Y.D.: Dual-path convolutional image-text embeddings with instance loss. ACM Transactions on Multimedia Computing, Communications, and Applications 16
2020
Earlier work this paper cites.
Zhu, F., Zhu, Y., Chang, X., Liang, X.: Vision-and-language navigation with self-supervised auxiliary reasoning tasks. In: CVPR (2020)
2020
Earlier work this paper cites.
Dai, M., Hu, J., Zhuang, J., Zheng, E.: A transformer-based feature segmentation and region alignment method for uav-view geo-localization. IEEE Transactions on Circuits and Systems for Video Technology 32
2021
Cited alongside, same era.
Hong, Y., Rodriguez-Opazo, C., Wu, Q., Gould, S.: Vln-bert: A recurrent vision-and-language bert for navigation. In: CVPR (2021)
2021
Cited alongside, same era.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: ICML. pp. 4904–4916 (2021)
2021
Cited alongside, same era.
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., Hoi, S.C.H.: Align before fuse: Vision and language representation learning with momentum distillation. In: NeurIPS. vol. 34, pp. 9694–9705 (2021)
2021
Cited alongside, same era.
Hämäläinen, P., Tavast, M., Kunnari, A.: Evaluating large language models in generating synthetic hci research data: a case study. In: CHI. pp. 1–19 (2023)
2023
Closest in time.
Hu, X., Hu, Y., Resch, B., Kersten, J.: Geographic information extraction from texts (geoext). In: ECCV. pp. 398–404 (2023)
2023
Closest in time.
Kuzman, T., Mozetic, I., Ljubešic, N.: Chatgpt: beginning of an end of manual linguistic data annotation. arXiv (2023)
2023
Closest in time.
Li, Y., Zhang, C., Yu, G., Wang, Z., Fu, B., Lin, G., Shen, C., Chen, L., Wei, Y.: Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue data. arXiv (2023)
2023
Closest in time.
Meng, Y., Michalski, M., Huang, J., Zhang, Y., Abdelzaher, T., Han, J.: Tuning language models as training data generators for augmentation-enhanced few-shot learning. In: ICML. pp. 24457–24477 (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: ICCV. pp. 10012–10022 (2021)
2021
Cited alongside, same era.
Pasquini, G., Arias, J.E.R., Schäfer, P., Busskamp, V.: Automated methods for cell type annotation on scrna-seq data. Computational and Structural Biotechnology Journal 19
2021
Cited alongside, same era.
Qi, Y., Pan, Z., Zhang, S., van den Hengel, A., Wu, Q.: Object-and-room informed sequential bert for vision-and-language navigation. In: ICCV (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763 (2021)
2021
Cited alongside, same era.
Rodrigues, R., Tani, M.: Are these from the same place? seeing the unseen in cross-view image geo-localization. In: WACV. pp. 3753–3761 (2021)
2021
Cited alongside, same era.
Wang, T., Zheng, Z., Yan, C., Zhang, J., Sun, Y., Zheng, B., Yang, Y.: Each part matters: Local patterns facilitate cross-view geo-localization. IEEE Transactions on Circuits and Systems for Video Technology 32
2021
Cited alongside, same era.
Yang, H., Lu, X., Zhu, Y.: Cross-view geo-localization with layer-to-layer transformer. In: NeurIPS. vol. 34, pp. 29009–29020 (2021)
2021
Cited alongside, same era.
Zhu, P., Wen, L., Du, D., Bian, X., Fan, H., Hu, Q., Ling, H.: Detection and tracking meet drones challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence 44
2021
Cited alongside, same era.
2023
Closest in time.
OpenAI: Gpt-4 technical report. arXiv (2023)
2023
Closest in time.
Pangakis, N., Wolken, S., Fasching, N.: Automated annotation with generative ai requires validation. arXiv (2023)
2023
Closest in time.
Shvetsova, N., Kukleva, A., Hong, X., Rupprecht, C., Schiele, B., Kuehne, H.: Howtocaption: Prompting llms to transform video annotations at scale. arXiv (2023)
2023
Closest in time.
Sun, B., Liu, G., Yuan, Y.: F3-net: Multiview scene matching for drone-based geo-localization. IEEE Transactions on Geoscience and Remote Sensing 61
2023
Closest in time.
Trivigno, G., Berton, G., Aragon, J., Caputo, B., Masone, C.: Divide&classify: Fine-grained classification for city-wide visual geo-localization. In: ICCV. pp. 11142–11152 (2023)
2023
Closest in time.
Wang, K., Fu, X., Huang, Y., Cao, C., Shi, G., Zha, Z.J.: Generalized uav object detection via frequency domain disentanglement. In: CVPR. pp. 1064–1073 (2023)
2023
Closest in time.
Wang, W., Lin, X., Feng, F., He, X., Chua, T.S.: Generative recommendation: Towards next-generation recommender paradigm. arXiv (2023)
2023
Closest in time.
Yang, S., Zhou, Y., Zheng, Z., Wang, Y., Zhu, L., Wu, Y.: Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. In: ACM MM. pp. 4492–4501 (2023)
2023
Closest in time.
Yu, W., Iter, D., Wang, S., Xu, Y., Ju, M., Sanyal, S., Zhu, C., Zeng, M., Jiang, M.: Generate rather than retrieve: Large language models are strong context generators. In: ICLR (2023)
2023
Closest in time.
Zhang, R., Li, Y., Ma, Y., Zhou, M., Zou, L.: Llmaaa: Making large language models as active annotators. In: EMNLP. pp. 13088–13103 (2023)
2023
Closest in time.
Zhang, X., Li, X., Sultani, W., Zhou, Y., Wshah, S.: Cross-view geo-localization via learning disentangled geometric layout correspondence. In: AAAI. vol. 37, pp. 3480–3488 (2023)
2023
Closest in time.
2023
Closest in time.
Chen, D., Chen, R., Zhang, S., Liu, Y., Wang, Y., Zhou, H., Zhang, Q., Zhou, P., Wan, Y., Sun, L.: Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. In: ICML (2024)
2024
Closest in time.
Chen, G.H., Chen, S., Zhang, R., Chen, J., Wu, X., Zhang, Z., Chen, Z., Li, J., Wan, X., Wang, B.: Allava: Harnessing gpt4v-synthesized data for a lite vision-language model. arXiv (2024)
2024
Closest in time.
Dhakal, A., Ahmad, A., Khanal, S., Sastry, S., Jacobs, N.: Sat2cap: Mapping fine-grained textual descriptions from satellite images. In: CVPR Workshops. pp. 533–542 (2024)
2024
Closest in time.
Ikezogwo, W., Seyfioglu, S., Ghezloo, F., Geva, D., Sheikh Mohammed, F., Anand, P.K., Krishna, R., Shapiro, L.: Quilt-1m: One million image-text pairs for histopathology. In: NeurIPS. vol. 36 (2024)
2024
Closest in time.
Liu, S., Hussain, A.S., Sun, C., Shan, Y.: Music understanding llama: Advancing text-to-music generation with question answering and captioning. In: ICASSP. pp. 286–290 (2024)
2024
Closest in time.
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al.: Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In: ECCV (2024)
2024
Closest in time.
Maaz, M., Rasheed, H., Khan, S., Khan, F.S.: Video-chatgpt: Towards detailed video understanding via large vision and language models. In: ACL (2024)
2024
Closest in time.
Wang, T., Zheng, Z., Sun, Y., Chua, T.S., Yang, Y., Yan, C.: Multiple-environment self-adaptive network for aerial-view geo-localization. Pattern Recognition (2024)
2024
Closest in time.
Yu, Y., Zhuang, Y., Zhang, J., Meng, Y., Ratner, A.J., Krishna, R., Shen, J., Zhang, C.: Large language model as attributed training data generator: A tale of diversity and bias. NeurIPS 36
2024
Closest in time.
Zhu, D., Chen, J., Shen, X., Li, X., Elhoseiny, M.: Minigpt-4: Enhancing vision-language understanding with advanced large language models. In: ICLR (2024)
2024
Closest in time.
Zhu, W., Hessel, J., Awadalla, A., Gadre, S.Y., Dodge, J., Fang, A., Yu, Y., Schmidt, L., Wang, W.Y., Choi, Y.: Multimodal c4: An open, billion-scale corpus of images interleaved with text. NeurIPS 36
2024
Closest in time.