Fetching the paper…
Reading the bibliography…
Seemingly simple natural language requests to a robot are generally underspecified, for example "Can you bring me the wireless mouse?" Flat images of candidate mice may not provide the discriminative information needed for "wireless." The world, and objects in it, are not flat images but complex 3D shapes.
Wordnet: A lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Earlier work this paper cites.
Going beyond text: A hybrid image-text approach for measuring word relatedness
C. W. Leong and R. Mihalcea · 2011
Earlier work this paper cites.
First-person vision
T. Kanade and M. Hebert · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
ShapeNet: An Information-Rich 3D Model Repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, J. Xiao, L. Yi, and F. Yu · 2015
Earlier work this paper cites.
Semantically-enriched 3d models for common-sense knowledge
A. X. Chang, M. Savva, and P. Hanrahan · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning multi-modal grounded linguistic semantics by playing “I spy”
J. Thomason, J. Sinapov, M. Svetlik, P. Stone, and R. J. Mooney · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Dynamic motion learning for multi-dof flexible-joint robots using active–passive motor babbling through deep learning
K. Takahashi, T. Ogata, J. Nakanishi, G. Cheng, and S. Sugano · 2017
Earlier work this paper cites.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
H. De Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville · 2017
Earlier work this paper cites.
Generalized grounding graphs: A probabilistic framework for understanding grounded commands
T. Kollar, S. Tellex, M. R. Walter, A. S. Huang, A. Bachrach, S. Hemachandra, E. Brunskill, A. Banerjee, D. Roy, S. Teller, and N. Roy · 2017
Earlier work this paper cites.
Yale-cmu-berkeley dataset for robotic manipulation research
B. Calli, A. Singh, J. Bruce, A. Walsman, K. Konolige, S. Srinivasa, P. Abbeel, and A. M. Dollar · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Jointly improving parsing and perception for natural language commands through human-robot dialog
J. Thomason, A. Padmakumar, J. Sinapov, N. Walker, Y. Jiang, H. Yedidsion, J. Hart, P. Stone, and R. J. Mooney · 2020
Later among the works it cites.
Grounding language in play
C. Lynch and P. Sermanet · 2020
Later among the works it cites.
Experience grounds language
Y. Bisk, A. Holtzman, J. Thomason, J. Andreas, Y. Bengio, J. Chai, M. Lapata, A. Lazaridou, J. May, A. Nisnevich, N. Pinto, and J. Turian · 2020
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee · 2020
Later among the works it cites.
Imitating interactive intelligence
J. Abramson, A. Ahuja, A. Brussee, F. Carnevale, M. Cassin, S. Clark, A. Dudzik, P. Georgiev, A. Guy, T. Harley, F. Hill, A. Hung, Z. Kenton, J. Landon, T. Lillicrap, K. Mathewson, A. Muldal, A. Santoro, N. Savinov, V. Varma, G. Wayne, N. Wong, C. Yan, and R. Zhu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VisualBERT: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Cited alongside, same era.
6-dof graspnet: Variational grasp generation for object manipulation
A. Mousavian, C. Eppner, and D. Fox · 2019
Cited alongside, same era.
Grounding language attributes to objects using bayesian eigenobjects
V. Cohen, B. Burchfiel, T. Nguyen, N. Gopalan, S. Tellex, and G. Konidaris · 2019
Cited alongside, same era.
ShapeGlot: Learning language for shape differentiation
P. Achlioptas, J. Fan, R. X. Hawkins, N. D. Goodman, and L. J. Guibas · 2019
Cited alongside, same era.
Improving robot success detection using static object data
R. Scalise, J. Thomason, Y. Bisk, and S. Srinivasa · 2019
Cited alongside, same era.
Prospection: Interpretable plans from language by predicting the future
C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Fox · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Later among the works it cites.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Later among the works it cites.
A recurrent vision-and-language BERT for navigation
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould · 2020
Later among the works it cites.
MOCA: A modular object-centric approach for interactive instruction following
K. P. Singh, S. Bhambri, B. Kim, R. Mottaghi, and J. Choi · 2020
Later among the works it cites.
INGRESS: Interactive visual grounding of referring expressions
M. Shridhar, D. Mittal, and D. Hsu · 2020
Later among the works it cites.
Goal-aware prediction: Learning to model what matters
S. Nair, S. Savarese, and C. Finn · 2020
Later among the works it cites.
ACRONYM: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2020
Later among the works it cites.
Learning rgb-d feature embeddings for unseen object instance segmentation
Y. Xiang, C. Xie, A. Mousavian, and D. Fox · 2020
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Closest in time.
Planning multimodal exploratory actions for online robot attribute learning
X. Zhang, J. Sinapov, and S. Zhang · 2021
Closest in time.
Pathdreamer: A world model for indoor navigation
J. Y. Koh, H. Lee, Y. Yang, J. Baldridge, and P. Anderson · 2021
Closest in time.
VinVL: Making visual representations matter in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Closest in time.