Fetching the paper…
Reading the bibliography…
The human language is one of the most natural interfaces for humans to interact with robots.
An efficient k-means clustering algorithm: Analysis and implementation
T. Kanungo, D. M. Mount, N. S. Netanyahu, C. D. Piatko, R. Silverman, and A. Y. Wu · 2002
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
K. Toutanova, D. Klein, C. D. Manning, and Y. Singer · 2003
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
D. G. Lowe · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Semantic place classification of indoor environments with mobile robots using boosting
A. Rottmann, Ó. M. Mozos, C. Stachniss, and W. Burgard · 2005
Earlier work this paper cites.
Surf: Speeded up robust features
H. Bay, T. Tuytelaars, and L. Van Gool · 2006
Earlier work this paper cites.
A game-theoretic approach to generating spatial descriptions
D. Golland, P. Liang, and D. Klein · 2010
Earlier work this paper cites.
Grounding spatial language for video search
S. Tellex, T. Kollar, G. Shaw, N. Roy, and D. Roy · 2010
Earlier work this paper cites.
Autonomous semantic mapping for robots performing everyday manipulation tasks in kitchen environments
N. Blodow, L. C. Goron, Z.-C. Marton, D. Pangercic, T. Rühr, M. Tenorth, and M. Beetz · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A Joint Model of Language and Perception for Grounded Attribute Learning
C. Matuszek, N. FitzGerald, L. Zettlemoyer, L. Bo, and D. Fox · 2012
Earlier work this paper cites.
Semantic object maps for robotic housework-representation, acquisition and use
D. Pangercic, B. Pitzer, M. Tenorth, and M. Beetz · 2012
Earlier work this paper cites.
Toward information theoretic human-robot dialog
S. Tellex, P. Thaker, R. Deits, T. Kollar, and N. Roy · 2012
Earlier work this paper cites.
Learning distributions over logical forms for referring expression generation
N. FitzGerald, Y. Artzi, and L. S. Zettlemoyer · 2013
Earlier work this paper cites.
Grounding spatial relations for human-robot interaction
S. Guadarrama, L. Riano, D. Golland, D. Go, Y. Jia, D. Klein, P. Abbeel, T. Darrell, et al · 2013
Cited alongside, same era.
Building semantic object maps from sparse and noisy 3d data
M. Gunther, T. Wiemann, S. Albrecht, and J. Hertzberg · 2013
Cited alongside, same era.
Multiscale combinatorial grouping
P. Arbeláez, J. Pont-Tuset, J. T. Barron, F. Marques, and J. Malik · 2014
Cited alongside, same era.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Lessons from the amazon picking challenge: Four aspects of building robotic systems
C. Eppner, S. Höfer, R. Jonschkowski, R. Martın-Martın, A. Sieverling, V. Wall, and O. Brock · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
Natural spatial description generation for human-robot interaction in indoor environments
Z. Huo and M. Skubic · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Deep learning for detecting robotic grasps
I. Lenz, H. Lee, and A. Saxena · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen · 2016
Later among the works it cites.
Spatial references and perspective in natural language instructions for collaborative manipulation
S. Li, R. Scalise, H. Admoni, S. Rosenthal, and S. S. Srinivasa · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Later among the works it cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Later among the works it cites.
Efficient grounding of abstract spatial concepts for natural language interaction with robot manipulators
R. Paul, J. Arkin, N. Roy, and T. Howard · 2016
Later among the works it cites.
Recognising the clothing categories from free-configuration using gaussian-process-based interactive perception
L. Sun, S. Rogers, G. Aragon-Camarasa, and J. P. Siebert · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.
Guesswhat?! visual object discovery through multi-modal dialogue
H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville · 2017
Closest in time.
A joint speaker-listener-reinforcer model for referring expressions
L. Yu, H. Tan, M. Bansal, and T. L. Berg · 2017
Closest in time.