Fetching the paper…
Reading the bibliography…
Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 1988–1997, 2016. [Online]. Available: https://api.semanticscholar.org/CorpusID:15458100
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
H. H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in Conference on Empirical Methods in Natural Language Processing , 2019
2019
Earlier work this paper cites.
K. Yang, O. Russakovsky, and J. Deng, “Spatialsense: An adversarially crowdsourced benchmark for spatial relation recognition,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 2051–2060, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Li, R. R. Selvaraju, A. D. Gotmare, S. R. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in Neural Information Processing Systems , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1780–1790
2021
Cited alongside, same era.
B. Cheng, A. G. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” in Neural Information Processing Systems , 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:235829267
2021
Cited alongside, same era.
H. Ha and S. Song, “Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models,” 2022
2022
Cited alongside, same era.
Y. Mo, H. Zhang, and T. Kong, “Towards open-world interactive disambiguation for robotic grasping,” in CoRL 2022 Workshop on Learning, Perception, and Abstraction for Long-Horizon Planning , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
A. Kamath, J. Hessel, and K.-W. Chang, “What’s “up” with vision-language models? investigating their struggle with spatial reasoning,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 9161–9175. [Online]. Available: https://aclanthology.org/2023.emnlp-main.568
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
Cited alongside, same era.
F. Liu, G. E. T. Emerson, and N. Collier, “Visual spatial reasoning,” Transactions of the Association for Computational Linguistics , vol. 11, pp. 635–651, 2022
2022
Cited alongside, same era.
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra, “Detecting twenty-thousand classes using image-level supervision,” in ECCV , 2022
2022
Cited alongside, same era.
S. Subramanian, W. Merrill, T. Darrell, M. Gardner, S. Singh, and A. Rohrbach, “Reclip: A strong zero-shot baseline for referring expression comprehension,” in Annual Meeting of the Association for Computational Linguistics , 2022
2022
Cited alongside, same era.
T. Gupta, A. Kamath, A. Kembhavi, and D. Hoiem, “Towards general purpose vision systems: An end-to-end task-agnostic vision-language architecture,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 16 399–16 409
2022
Cited alongside, same era.
B. Jiang and C. J. Taylor, “Hierarchical relationships: A new perspective to enhance scene graph generation,” 2023
2023
Cited alongside, same era.
N. Nejatishahidin, W. Hutchcroft, M. Narayana, I. Boyadzhiev, Y. Li, N. Khosravan, J. Košecká, and S. B. Kang, “Graph-covis: Gnn-based multi-view panorama global pose estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6458–6467
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. H. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence, “Palm-e: An embodied multimodal language model,” in International Conference on Machine Learning , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:257364842
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. B. Girshick, “Segment anything,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 3992–4003, 2023
2023
Later among the works it cites.
Y. Li, N. Rajabi, S. Shrestha, M. A. Reza, and J. Kosecka, “Labeling indoor scenes with fusion of out-of-the-box perception models,” 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) , pp. 570–579, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:269203025
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia, “Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2024, pp. 14 455–14 465
2024
Closest in time.