Fetching the paper…
Reading the bibliography…
Controlling robots to perform tasks via natural language is one of the most challenging topics in human-robot interaction.
Grounding in communication
Herbert H Clark and Susan E Brennan · 1991
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Semantic object maps for robotic housework-representation, acquisition and use
Dejan Pangercic, Benjamin Pitzer, Moritz Tenorth, and Michael Beetz · 2012
Earlier work this paper cites.
Learning to place new objects in a scene
Yun Jiang, Marcus Lim, Changxi Zheng, and Ashutosh Saxena · 2012
Earlier work this paper cites.
Grounding spatial relations for human-robot interaction
Sergio Guadarrama, Lorenzo Riano, Dave Golland, Daniel Go, Yangqing Jia, Dan Klein, Pieter Abbeel, and Trevor Darrell · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Efficient grounding of abstract spatial concepts for natural language interaction with robot manipulators
Rohan Paul, Jacob Arkin, Nicholas Roy, and Thomas M Howard · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
High precision grasp pose detection in dense clutter
Marcus Gualtieri, Andreas Ten Pas, Kate Saenko, and Robert Platt · 2016
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Cited alongside, same era.
Organizing objects by predicting user preferences through collaborative filtering
Nichola Abdo, Cyrill Stachniss, Luciano Spinello, and Wolfram Burgard · 2016
Cited alongside, same era.
Metric learning for generalizing spatial relations to new objects
Oier Mees, Nichola Abdo, Mladen Mazuran, and Wolfram Burgard · 2017
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Later among the works it cites.
Object placement planning and optimization for robot manipulators
Joshua Alexander Haustein, Kaiyu Hang, Johannes A Stork, and Danica Kragic · 2019
Later among the works it cites.
Self-supervised 3d shape and viewpoint estimation from single images for robotics
Oier Mees, Maxim Tatarchenko, Thomas Brox, and Wolfram Burgard · 2019
Later among the works it cites.
6-dof graspnet: Variational grasp generation for object manipulation
Arsalan Mousavian, Clemens Eppner, and Dieter Fox · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
A joint speaker-listener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Cited alongside, same era.
Shape completion enabled robotic grasping
Jacob Varley, Chad DeChant, Adam Richardson, Joaquín Ruales, and Peter Allen · 2017
Cited alongside, same era.
Interactively picking real-world objects with unconstrained spoken language instructions
Jun Hatori, Yuta Kikuchi, Sosuke Kobayashi, Kuniyuki Takahashi, Yuta Tsuboi, Yuya Unno, Wilson Ko, and Jethro Tan · 2018
Cited alongside, same era.
Interactive visual grounding of referring expressions for human-robot interaction
Mohit Shridhar and David Hsu · 2018
Cited alongside, same era.
Learning object placements for relational instructions by hallucinating scene representations
Oier Mees, Alp Emek, Johan Vertens, and Wolfram Burgard · 2020
Later among the works it cites.
Adversarial skill networks: Unsupervised robot skill learning from videos
Oier Mees, Markus Merklinger, Gabriel Kalweit, and Wolfram Burgard · 2020
Later among the works it cites.
Hindsight for foresight: Unsupervised structured dynamics models from physical interaction
Iman Nematollahi, Oier Mees, Lukas Hermann, and Wolfram Burgard · 2020
Later among the works it cites.
Corey Lynch and Pierre Sermanet · 2020
Later among the works it cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2020
Later among the works it cites.