Fetching the paper…
Reading the bibliography…
This paper presents INVIGORATE, a robot system that interacts with human through natural language and grasps a specified object in clutter.
arXiv preprint arXiv:1905.12197
Cai P, Luo Y, Saxena A, Hsu D and Lee WS (2019) Lets-drive: Driving in a crowd by learning from tree search · 1905
Earlier work this paper cites.
arXiv preprint arXiv:1908.02265
Lu J, Batra D, Parikh D and Lee S (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks · 1908
Earlier work this paper cites.
Neural computation 9(8): 1735–1780
Hochreiter S and Schmidhuber J (1997) Long short-term memory · 1997
Earlier work this paper cites.
Artificial intelligence 101(1-2): 99–134
Kaelbling LP, Littman ML and Cassandra AR (1998) Planning and acting in partially observable stochastic domains · 1998
Earlier work this paper cites.
In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics . pp. 311–318
Papineni K, Roukos S, Ward T and Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation · 2002
Earlier work this paper cites.
In: Text summarization branches out . pp. 74–81
Lin CY (2004) Rouge: A package for automatic evaluation of summaries · 2004
Earlier work this paper cites.
In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . pp. 65–72
Banerjee S and Lavie A (2005) Meteor: An automatic metric for mt evaluation with improved correlation with human judgments · 2005
Earlier work this paper cites.
In: Proceedings of the 1st ACM SIGCHI/SIGART conference on Human-robot interaction . pp. 282–289
Kruijff GJM, Zender H, Jensfelt P and Christensen HI (2006) Clarification dialogues in human-augmented mapping · 2006
Earlier work this paper cites.
In: Proceedings of the 25th international conference on Machine learning . pp. 240–247
Diuk C, Cohen A and Littman ML (2008) An object-oriented representation for efficient reinforcement learning · 2008
Earlier work this paper cites.
Foundations and Trends® in Human–Computer Interaction 1(3): 203–275
Goodrich MA, Schultz AC et al. (2008) Human–robot interaction: A survey · 2008
Earlier work this paper cites.
The knowledge engineering review 25(1): 1–25
Foulds J and Frank E (2010) A review of multi-instance learning assumptions · 2010
Earlier work this paper cites.
In: AAMAS , volume 10. pp. 915–922
Rosenthal S, Biswas J and Veloso MM (2010) An effective personal mobile robot agent through symbiotic human-robot interaction · 2010
Earlier work this paper cites.
IEEE Transactions on Robotics 30(2): 289–309
Bohg J, Morales A, Asfour T and Kragic D (2013) Data-driven grasp synthesis—a survey · 2013
Earlier work this paper cites.
Journal of Human-Robot Interaction 2(2): 58–79
Deits R, Tellex S, Thaker P, Simeonov D, Kollar T and Roy N (2013) Clarifying commands with information-theoretic human-robot dialog · 2013
Earlier work this paper cites.
In: Robotics: science and systems
Guadarrama S, Rodner E, Saenko K, Zhang N, Farrell R, Donahue J and Darrell T (2014) Open-vocabulary object retrieval · 2014
Earlier work this paper cites.
In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . pp. 787–798
Kazemzadeh S, Ordonez V, Matten M and Berg T (2014) Referitgame: Referring to objects in photographs of natural scenes · 2014
Earlier work this paper cites.
In: European conference on computer vision . Springer, pp. 740–755
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P and Zitnick CL (2014) Microsoft coco: Common objects in context · 2014
Earlier work this paper cites.
Tellex S, Knepper R, Li A, Rus D and Roy N (2014) Asking for help using inverse semantics
2014
Earlier work this paper cites.
In: 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 5115–5121
Hemachandra S and Walter MR (2015) Information-theoretic dialog to improve spatial-semantic representations · 2015
Earlier work this paper cites.
The International Journal of Robotics Research 34(4-5): 705–724
Lenz I, Lee H and Saxena A (2015) Deep learning for detecting robotic grasps · 2015
Earlier work this paper cites.
In: 2015 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 1316–1322
Redmon J and Angelova A (2015) Real-time grasp detection using convolutional neural networks · 2015
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 4566–4575
Vedantam R, Lawrence Zitnick C and Parikh D (2015) Cider: Consensus-based image description evaluation · 2015
Earlier work this paper cites.
In: European conference on computer vision . Springer, pp. 382–398
Anderson P, Fernando B, Johnson M and Gould S (2016) Spice: Semantic propositional image caption evaluation · 2016
Earlier work this paper cites.
Journal of Artificial Intelligence Research 55: 409–442
Bernardi R, Cakici R, Elliott D, Erdem A, Erdem E, Ikizler-Cinbis N, Keller F, Muscat A and Plank B (2016) Automatic description generation from images: A survey of models, datasets, and evaluation measures · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 4565–4574
Johnson J, Karpathy A and Fei-Fei L (2016) Densecap: Fully convolutional localization networks for dense captioning · 2016
Earlier work this paper cites.
In: 2016 25th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IEEE, pp. 44–51
Li S, Scalise R, Admoni H, Rosenthal S and Srinivasa SS (2016) Spatial references and perspective in natural language instructions for collaborative manipulation · 2016
Cited alongside, same era.
In: European Conference on Computer Vision . Springer, pp. 792–807
Nagaraja VK, Morariu VI and Davis LS (2016) Modeling context between objects for referring expression understanding · 2016
Cited alongside, same era.
nature 529(7587): 484–489
Silver D, Huang A, Maddison CJ, Guez A, Sifre L, Van Den Driessche G, Schrittwieser J, Antonoglou I, Panneershelvam V, Lanctot M et al. (2016) Mastering the game of go with deep neural networks and tree search · 2016
Cited alongside, same era.
In: Proceedings of the 30th International Conference on Neural Information Processing Systems . pp. 2154–2162
Tamar A, Wu Y, Thomas G, Levine S and Abbeel P (2016) Value iteration networks · 2016
Cited alongside, same era.
In: Conference on Robot Learning . PMLR, pp. 119–132
Jang E, Vijayanarasimhan S, Pastor P, Ibarz J and Levine S (2017) End-to-end learning of semantic grasping · 2017
Science Robotics 4(26): eaau4984
Mahler J, Matl M, Satish V, Danielczuk M, DeRose B, McKinley S and Goldberg K (2019) Learning ambidextrous robot grasping policies · 2019
Later among the works it cites.
In: The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, pp. 3459–3467
Vaicenavicius J, Widmann D, Andersson C, Lindsten F, Roll J and Schön T (2019) Evaluating model calibration in classification · 2019
Later among the works it cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 7194–7200
Wandzel A, Oh Y, Fishman M, Kumar N, Wong LL and Tellex S (2019) Multi-object search using object-oriented pomdps · 2019
Later among the works it cites.
The International Journal of Robotics Research : 0278364919868017
Zeng A, Song S, Yu KT, Donlon E, Hogan FR, Bauza M, Ma D, Taylor O, Liu M, Romo E et al. (2019) Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching · 2019
Later among the works it cites.
In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 6435–6442
Zhang H, Lan X, Bai S, Wan L, Yang C and Zheng N (2019) A multi-task convolutional neural network for autonomous robotic grasping in object stacking scenes · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: Advances in Neural Information Processing Systems , volume 30. Curran Associates, Inc
Karkus P, Hsu D and Lee WS (2017) Qmdp-net: Deep learning for planning under partial observability · 2017
Cited alongside, same era.
International journal of computer vision 123(1): 32–73
Krishna R, Zhu Y, Groth O, Johnson J, Hata K, Kravitz J, Chen S, Kalantidis Y, Li LJ, Shamma DA et al. (2017) Visual genome: Connecting language and vision using crowdsourced dense image annotations · 2017
Cited alongside, same era.
arXiv preprint arXiv:1703.09312
Mahler J, Liang J, Niyaz S, Laskey M, Doan R, Liu X, Ojea JA and Goldberg K (2017) Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics · 2017
Cited alongside, same era.
Advances in neural information processing systems 30
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł and Polosukhin I (2017) Attention is all you need · 2017
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 2193–2202
Yang L, Tang K, Yang J and Li LJ (2017) Dense captioning with joint inference and visual context · 2017
Cited alongside, same era.
Advances in neural information processing systems 31
Amos B, Jimenez I, Sacks J, Boots B and Kolter JZ (2018) Differentiable mpc for end-to-end planning and control · 2018
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 6154–6162
Cai Z and Vasconcelos N (2018) Cascade r-cnn: Delving into high quality object detection · 2018
Cited alongside, same era.
Later among the works it cites.
In: European Conference on Computer Vision . Springer, pp. 104–120
Chen YC, Li L, Yu L, El Kholy A, Ahmed F, Gan Z, Cheng Y and Liu J (2020) Uniter: Universal image-text representation learning · 2020
Later among the works it cites.
In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 8408–8414
Kurenkov A, Taglic J, Kulkarni R, Dominguez-Kuhne M, Garg A, Martín-Martín R and Savarese S (2020) Visuomotor mechanical search: Learning to retrieve target objects in clutter · 2020
Later among the works it cites.
In: 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 6232–6238
Murali A, Mousavian A, Eppner C, Paxton C and Fox D (2020) 6-dof grasping for target-driven object manipulation in clutter · 2020
Later among the works it cites.
IEEE Transactions on Multimedia
Qiao Y, Deng C and Wu Q (2020) Referring expression comprehension: A survey of methods and datasets · 2020
Later among the works it cites.
In: International Conference on Machine Learning . PMLR, pp. 8583–8592
Sekar R, Rybkin O, Daniilidis K, Abbeel P, Hafner D and Pathak D (2020) Planning to explore via self-supervised world models · 2020
Later among the works it cites.
The International Journal of Robotics Research 39(2-3): 217–232
Shridhar M, Mittal D and Hsu D (2020) Ingress: Interactive visual grounding of referring expressions · 2020
Later among the works it cites.
IEEE Robotics and Automation Letters 5(2): 2232–2239
Yang Y, Liang H and Choi C (2020) A deep learning approach to grasping the invisible · 2020
Later among the works it cites.
Artificial Intelligence 297: 103500
Arora S and Doshi P (2021) A survey of inverse reinforcement learning: Challenges, methods and progress · 2021
Closest in time.
arXiv preprint arXiv:2104.00492
Chen Y, Xu R, Lin Y and Vela PA (2021) A joint network for grasp detection conditioned on natural language commands · 2021
Closest in time.
In: Proceedings of the IEEE/CVF International Conference on Computer Vision . pp. 1769–1779
Deng J, Yang Z, Chen T, Zhou W and Li H (2021) Transvg: End-to-end visual grounding with transformers · 2021
Closest in time.
In: Proceedings of the IEEE/CVF International Conference on Computer Vision . pp. 1780–1790
Kamath A, Singh M, LeCun Y, Synnaeve G, Misra I and Carion N (2021) Mdetr-modulated detection for end-to-end multi-modal understanding · 2021
Closest in time.
arXiv preprint arXiv:2102.11506
Katiyar S and Borgohain SK (2021) Comparative evaluation of cnn architectures for image caption generation · 2021
Closest in time.
Advances in Neural Information Processing Systems 34: 19652–19664
Li M and Sigal L (2021) Referring transformer: A one-step approach to multi-task visual grounding · 2021
Closest in time.
In: International Symposium on Experimental Robotics (ISER)
Mees O and Burgard W (2021) Composing pick-and-place tasks by grounding language · 2021
Closest in time.
In: 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 4209–4215
Okada M and Taniguchi T (2021) Dreaming: Model-based reinforcement learning by latent imagination without reconstruction · 2021
Closest in time.
In: Proceedings of Robotics: Science and Systems . Virtual
Zhang H, Lu Y, Yu C, Hsu D, Lan X and Zheng N (2021) INVIGORATE: Interactive Visual Grounding and Grasping in Clutter · 2021
Closest in time.
In: Proceedings of the 39th international conference on Machine learning
Zeng Y, Zhang X and Li H (2022) Multi-grained vision language pre-training: Aligning texts with visual concepts · 2022
Closest in time.
Foundations and Trends® in Machine Learning 16(1): 1–118
Moerland TM, Broekens J, Plaat A, Jonker CM et al. (2023) Model-based reinforcement learning: A survey · 2023
Closest in time.
In: 2016 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 2038–2043
Guo D, Kong T, Sun F and Liu H (2016) Object discovery and grasp detection with a shared convolutional neural network · 2043
Closest in time.