Fetching the paper…
Reading the bibliography…
Most visual grounding solutions primarily focus on realistic images.
Cross-modal self-attention network for referring image segmentation
Ye, L., Rochan, M., Liu, Z., Wang, Y., 2019 · 1904
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M., Le, Q.V., 2019 · 1905
Earlier work this paper cites.
A fast and accurate one-stage approach to visual grounding
Yang, Z., Gong, B., Wang, L., Huang, W., Yu, D., Luo, J., 2019 · 1908
Earlier work this paper cites.
Iou loss for 2d/3d object detection
Zhou, D., Fang, J., Song, X., Guan, C., Yin, J., Dai, Y., Yang, R., 2019 · 1908
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L., 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J., 2019 · 1910
Earlier work this paper cites.
Large scale learning of general visual representations for transfer
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., Houlsby, N., 2019 · 1912
Earlier work this paper cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Zhang, J., Zhao, Y., Saleh, M., Liu, P.J., 2020 · 1912
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D., 2020 · 2005
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S., 2020 · 2005
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., Baldridge, J., 2020 · 2005
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J., 2021 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N., 2020 · 2010
Earlier work this paper cites.
Deformable DETR: deformable transformers for end-to-end object detection
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J., 2020 · 2010
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J., 2015 · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A.L., Murphy, K., 2015 · 2015
Earlier work this paper cites.
Deep interactive object selection
Xu, N., Price, B.L., Cohen, S., Yang, J., Huang, T.S., 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A.C., Berg, T.L., 2016 · 2016
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and VQA
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2017 · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents, in: Precup, D., Teh, Y.W. (Eds.), Proceedings of the 34th International Conference on Machine Learning, PMLR. pp. 3135–3144
Shi, T., Karpathy, A., Fan, L., Hernandez, J., Liang, P., 2017 · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Gur, I., Rückert, U., Faust, A., Hakkani-Tür, D., 2018 · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E.Z., Guu, K., Pasupat, P., Shi, T., Liang, P., 2018 · 2018
Earlier work this paper cites.
Beit: BERT pre-training of image transformers
Bao, H., Dong, L., Wei, F., 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D.A., Adeli, E., Altman, R.B., Arora, S., von Arx, S., Bernstein, M.S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N.S., Chen, A.S., Creel, K., Davis, J.Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N.D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D.E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P.W., Krass, M.S., Krishna, R., Kuditipudi, R., et al., 2021 · 2021
Earlier work this paper cites.
You only look at one sequence: Rethinking transformer in vision through object detection
Fang, Y., Liao, B., Wang, X., Fang, J., Qi, J., Wu, R., Niu, J., Liu, W., 2021 · 2021
Earlier work this paper cites.
Levit: a vision transformer in convnet’s clothing for faster inference
Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., Jégou, H., Douze, M., 2021 · 2021
Earlier work this paper cites.
Perceiver IO: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Hénaff, O.J., Botvinick, M.M., Zisserman, A., Vinyals, O., Carreira, J., 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y., Parekh, Z., Pham, H., Le, Q.V., Sung, Y., Li, Z., Duerig, T., 2021 · 2021
Cited alongside, same era.
Swin transformer V2: scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., Wei, F., Guo, B., 2021 · 2021
Cited alongside, same era.
Conditional DETR for fast training convergence
Meng, D., Chen, X., Fan, Z., Zeng, G., Li, H., Yuan, Y., Sun, L., Wang, J., 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I., 2021 · 2021
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., Lee, Y.J., 2023 · 2023
Later among the works it cites.
Fine-tuning large language models for adaptive machine translation
Moslem, Y., Haque, R., Way, A., 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Wei, F., 2023 · 2023
Later among the works it cites.
Real-time flying object detection with yolov8
Reis, D., Kupec, J., Hong, J., Daoudi, A., 2023 · 2023
Later among the works it cites.
From pixels to ui actions: Learning to follow instructions via graphical user interfaces, in: Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 34354–34370
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I., 2021 · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2021 · 2021
Cited alongside, same era.
Grounding natural language instructions: Can large language models capture spatial information?
Rozanova, J., Ferreira, D., Dubba, K., Cheng, W., Zhang, D., Freitas, A., 2021 · 2021
Cited alongside, same era.
Pubtables-1m: Towards comprehensive table extraction from unstructured documents
Smock, B., Pesala, R., Abraham, R., 2021 · 2021
Cited alongside, same era.
LAVT: language-aware vision transformer for referring image segmentation
Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., Torr, P.H.S., 2021 · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., Liu, C., Liu, M., Liu, Z., Lu, Y., Shi, Y., Wang, L., Wang, J., Xiao, B., Xiao, Z., Yang, J., Zeng, M., Zhou, L., Zhang, P., 2021 · 2021
Cited alongside, same era.
Bytetrack: Multi-object tracking by associating every detection box
Zhang, Y., Sun, P., Jiang, Y., Yu, D., Yuan, Z., Luo, P., Liu, W., Wang, X., 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., Simonyan, K., 2022 · 2022
Cited alongside, same era.
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., Toutanova, K.N., 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., Lample, G., 2023 · 2023
Later among the works it cites.
Ugif: Ui grounded instruction following
Venkatesh, S.G., Talukdar, P., Narayanan, S., 2023 · 2023
Later among the works it cites.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., Dai, J., 2023 · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., Huang, F., 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., Elhoseiny, M., 2023 · 2023
Later among the works it cites.
Segment everything everywhere all at once
Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J., Wang, L., Gao, J., Lee, Y.J., 2023 · 2023
Later among the works it cites.
MiniGPT-v2: Large language model as a unified interface for vision-language multi-task learning
Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., Elhoseiny, M., 2024 · 2024
Closest in time.
Seeclick: Harnessing gui grounding for advanced visual gui agents
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., Wu, Z., 2024 · 2024
Closest in time.
Assistgui: Task-oriented desktop graphical user interface automation
Gao, D., Ji, L., Bai, Z., Ouyang, M., Li, P., Mao, D., Wu, Q., Zhang, W., Wang, P., Guo, X., Wang, H., Zhou, L., Shou, M.Z., 2024 · 2024
Closest in time.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., Faust, A., 2024 · 2024
Closest in time.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y., Lee, Y.J., 2024 · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.W., Galley, M., Gao, J., 2024 · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.Y., Li, S.W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., Bojanowski, P., 2024 · 2024
Closest in time.
Visual grounding for user interfaces, in: Yang, Y., Davani, A., Sil, A., Kumar, A. (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), Association for Computational Linguistics, Mexico City, Mexico. pp. 97–107
Qian, Y., Lu, Y., Hauptmann, A., Riva, O., 2024 · 2024
Closest in time.
Corex: Pushing the boundaries of complex reasoning through multi-model collaboration
Sun, Q., Yin, Z., Li, X., Wu, Z., Qiu, X., Kong, L., 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V.I., Burnell, R., Bai, L., Gulati, A., et al., 2024 · 2024
Closest in time.
Contrastive decoding reduces hallucinations in large multilingual machine translation models, in: Graham, Y., Purver, M. (Eds.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, St. Julian’s, Malta. pp. 2526–2539
Waldendorf, J., Haddow, B., Birch, A., 2024 · 2024
Closest in time.
Symbol-llm: Towards foundational symbol-centric interface for large language models
Xu, F., Wu, Z., Sun, Q., Ren, S., Yuan, F., Yuan, S., Lin, Q., Qiao, Y., Liu, J., 2024 · 2024
Closest in time.
You only look at screens: Multimodal chain-of-action agents
Zhang, Z., Zhang, A., 2024 · 2024
Closest in time.
Gui-world: A video benchmark and dataset for multimodal gui-oriented understanding
Chen, D., Huang, Y., Wu, S., Tang, J., Chen, L., Bai, Y., He, Z., Wang, C., Zhou, H., Li, Y., Zhou, T., Yu, Y., Gao, C., Zhang, Q., Gui, Y., Li, Z., Wan, Y., Zhou, P., Gao, J., Sun, L., 2025 · 2025
Closest in time.
Winclick: Gui grounding with multimodal large language models
Hui, Z., Li, Y., zhao, D., Chen, T., Banbury, C., Koishida, K., 2025 · 2025
Closest in time.
Screenspot-pro: Gui grounding for professional high-resolution computer use
Li, K., Meng, Z., Lin, H., Luo, Z., Tian, Y., Ma, J., Huang, Z., Chua, T.S., 2025 · 2025
Closest in time.
Deskvision: Large scale desktop region captioning for advanced gui agents
Xu, Y., Yang, L., Chen, H., Wang, H., Chen, Z., Tang, Y., 2025 · 2025
Closest in time.