Fetching the paper…
Reading the bibliography…
For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O.-M. Camburu, A. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
Embodied question answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. V. Hengel · 2018
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson, A. X. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, and A. Zamir · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. L. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra · 2019
Cited alongside, same era.
Self-monitoring navigation agent via auxiliary progress estimation
C.-Y. Ma, J. Lu, Z. Wu, G. Al-Regib, Z. Kira, R. Socher, and C. Xiong · 2019
Cited alongside, same era.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
L. Ke, X. Li, Y. Bisk, A. Holtzman, Z. Gan, J. Liu, J. Gao, Y. Choi, and S. Srinivasa · 2019
Cited alongside, same era.
Conditional driving from natural language instructions
J. Roh, C. Paxton, A. Pronobis, A. Farhadi, and D. Fox · 2019
Topological planning with transformers for vision-and-language navigation
K. Chen, J. K. Chen, J. Chuang, M. V’azquez, and S. Savarese · 2020
Later among the works it cites.
A recurrent vision-and-language bert for navigation
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould · 2020
Later among the works it cites.
Sim-to-real transfer for vision-and-language navigation
P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee · 2020
Later among the works it cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee · 2020
Later among the works it cites.
Few-shot object grounding and mapping for natural language robot instruction following
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamic graph attention for referring expression comprehension
S. Yang, G. Li, and Y. Yu · 2019
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation, 2019
L. Ye, M. Rochan, Z. Liu, and Y. Wang · 2019
Cited alongside, same era.
Flair: An easy-to-use framework for state-of-the-art nlp
A. Akbik, T. Bergmann, D. Blythe, K. Rasul, S. Schweter, and R. Vollgraf · 2019
Cited alongside, same era.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas · 2020
Cited alongside, same era.
CLEVRER: collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2020
Cited alongside, same era.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
R. Girdhar and D. Ramanan · 2020
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Cited alongside, same era.
V. Blukis, R. A. Knepper, and Y. Artzi · 2020
Later among the works it cites.
Learning cross-modal context graph for visual grounding
Y. Liu, B. Wan, X.-D. Zhu, and X. He · 2020
Later among the works it cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Mdetr–modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Closest in time.
Visitron: Visual semantics-aligned interactively trained object-navigator
A. Shrivastava, K. Gopalakrishnan, Y. Liu, R. Piramuthu, G. Tur, D. Parikh, and D. Hakkani-Tur · 2021
Closest in time.
Hierarchical cross-modal agent for robotics vision-and-language navigation
M. Z. Irshad, C.-Y. Ma, and Z. Kira · 2021
Closest in time.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring, 2021
Z. Yuan, X. Yan, Y. Liao, R. Zhang, Z. Li, and S. Cui · 2021
Closest in time.
Free-form description guided 3d visual graph network for object grounding in point cloud, 2021
M. Feng, Z. Li, Q. Li, L. Zhang, X. Zhang, G. Zhu, H. Zhang, Y. Wang, and A. Mian · 2021
Closest in time.
SAT: 2d semantics assisted training for 3d visual grounding
Z. Yang, S. Zhang, L. Wang, and J. Luo · 2021
Closest in time.