Fetching the paper…
Reading the bibliography…
Humans are capable of completing a range of challenging manipulation tasks that require reasoning jointly over modalities such as vision, touch, and sound.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Social interaction of humanoid robot based on audio-visual tracking
Hiroshi Okuno, Kazuhiro Nakadai, and Hiroaki Kitano · 2002
Earlier work this paper cites.
Surveillance robot utilizing video and audio information
Xinyu Wu, Haitao Gong, Pei Chen, Zhi Zhong, and Yangsheng Xu · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell · 2011
Earlier work this paper cites.
Occlusion-aware reconstruction and manipulation of 3d articulated objects
Xiaoxia Huang, Ian Walker, and Stan Birchfield · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stable reinforcement learning with autoencoders for tactile and visual data
Herke Hoof, Nutan Chen, Maximilian Karl, Patrick van der Smagt, and Jan Peters · 2016
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh, and Honglak Lee · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh H. McDermott, Antonio Torralba, Edward H. Adelson, and William T. Freeman · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2016
Earlier work this paper cites.
Learning manipulation trajectories using recurrent neural networks
Rouhollah Rahmatizadeh, Pooya Abolghasemi, and Ladislau Bölöni · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
Jiakai Zhang and Kyunghyun Cho · 2016
Earlier work this paper cites.
Probabilistic articulated real-time tracking for robot manipulation
Cristina Garcia Cifuentes, Jan Issac, Manuel Wüthrich, Stefan Schaal, and Jeannette Bohg · 2017
Earlier work this paper cites.
Learning to navigate in complex environments
Piotr W. Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andy Ballard, Andrea Banino, Misha Denil, Ross Goroshin, L. Sifre, Koray Kavukcuoglu, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Vecerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Manfred Otto Heess, Thomas Rothörl, Thomas Lampe, and Martin A. Riedmiller · 2017
Earlier work this paper cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Earlier work this paper cites.
More than a feeling: Learning to grasp and regrasp using vision and touch
Roberto Calandra, Andrew Owens, Dinesh Jayaraman, Justin Lin, Wenzhen Yuan, Jitendra Malik, Edward H. Adelson, and Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning of active vision for manipulating objects under occlusions
Ricson Cheng, Arpit Agarwal, and Katerina Fragkiadaki · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, P. Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Cited alongside, same era.
Neural map: Structured memory for deep reinforcement learning
Emilio Parisotto and Ruslan Salakhutdinov · 2018
Cited alongside, same era.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Cited alongside, same era.
Occlusion-robust deformable object tracking without physics simulation
Cheng Chi and Dmitry Berenson · 2019
Cited alongside, same era.
Learning representations from audio-visual spatial alignment
Pedro Miguel Morgado, Ying Li, and Nuno Vasconcelos · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Çaglar Gülçehre, Siddhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matthew M. Botvinick, Nicolas Manfred Otto Heess, and Raia Hadsell · 2020
Later among the works it cites.
3d shape reconstruction from vision and touch
Edward Smith, Roberto Calandra, Adriana Romero, Georgia Gkioxari, David Meger, Jitendra Malik, and Michal Drozdzal · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robot learning for manipulation of granular materials using vision and sound
Samuel Clarke · 2019
Cited alongside, same era.
Mechanical search: Multi-step retrieval of a target object occluded by clutter
Michael Danielczuk, Andrey Kurenkov, Ashwin Balakrishna, Matthew Matl, David Wang, Roberto Martín-Martín, Animesh Garg, Silvio Savarese, and Ken Goldberg · 2019
Cited alongside, same era.
Scene memory transformer for embodied agents in long-horizon tasks
Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese · 2019
Cited alongside, same era.
Residual reinforcement learning for robot control
Tobias Johannink, Shikhar Bahl, Ashvin Nair, Jianlan Luo, Avinash Kumar, Matthias Loskyll, Juan Aparicio Ojea, Eugen Solowjow, and Sergey Levine · 2019
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
Michael Kelly, Chelsea Sidrane, K. Driggs-Campbell, and Mykel J. Kochenderfer · 2019
Cited alongside, same era.
Contextual reinforcement learning of visuo-tactile multi-fingered grasping policies
Visak C. V. Kumar, Tucker Hermans, Dieter Fox, Stan Birchfield, and Jonathan Tremblay · 2019
Cited alongside, same era.
Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks
Michelle A. Lee, Yuke Zhu, Krishna Parasuram Srinivasan, Parth Shah, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg · 2019
Cited alongside, same era.
Learning from interventions: Human-robot interaction as both explicit and implicit feedback
Jonathan Spencer, Sanjiban Choudhury, Matt Barnes, Matthew Schmittle, Mung Chiang, Peter Ramadge, and Siddhartha Srinivasa · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, Michael Laskin, and P. Abbeel · 2020
Later among the works it cites.
Event-driven visual-tactile sensing and learning for robots
Tasbolat Taunyazov, Weicong Sng, Hian Hian See, Brian Lim, Jethro Kuan, Abdul Fatir Ansari, Benjamin C. K. Tee, and Harold Soh · 2020
Later among the works it cites.
Robust, occlusion-aware pose estimation for objects grasped by adaptive hands
Bowen Wen, Chaitanya Mitash, Sruthi Soorian, Andrew Kimmel, Avishai Sintov, and Kostas E. Bekris · 2020
Later among the works it cites.
Learning 3d dynamic scene representations for robot manipulation
Zhenjia Xu, Zhanpeng He, Jiajun Wu, and Shuran Song · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2020
Later among the works it cites.
Residual reinforcement learning from demonstrations
Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, and Cordelia Schmid · 2021
Later among the works it cites.
Learning multimodal contact-rich skills from demonstrations without reward engineering
Mythra V. Balakuntala, Upinder Kaur, Xin Ma, Juan Pablo Wachs, and Richard M. Voyles · 2021
Later among the works it cites.
Occlusion-aware search for object retrieval in clutter
Wissam Bejjani, Wisdom C. Agboh, Mehmet Remzi Dogar, and Matteo Leonetti · 2021
Later among the works it cites.
Robotic grasping of fully-occluded objects using rf perception
Tara Boroushaki, Junshan Leng, Ian Clester, Alberto Rodriguez, and Fadel Adib · 2021
Later among the works it cites.
Structure from silence: Learning scene structure from ambient sound
Ziyang Chen, Xixi Hu, and Andrew Owens · 2021
Later among the works it cites.
Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations
Ruohan Gao, Yen-Yu Chang, Shivani Mall, Li Fei-Fei, and Jiajun Wu · 2021
Later among the works it cites.
Lazydagger: Reducing context switching in interactive imitation learning
Ryan Hoque, Ashwin Balakrishna, Carl Putterman, Michael Luo, Daniel S. Brown, Daniel Seita, Brijen Thananjeyan, Ellen R. Novoseller, and Ken Goldberg · 2021
Later among the works it cites.
Hideyuki Ichiwara, Hiroshi Ito, Kenjiro Yamamoto, Hiroki Mori, and Tetsuya Ogata · 2021
Later among the works it cites.
BC-z: Zero-shot task generalization with robotic imitation learning
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Aw-opt: Learning robotic skills with imitation and reinforcement at scale
Yao Lu, Karol Hausman, Yevgen Chebotar, Mengyuan Yan, Eric Jang, Alexander Herzog, Ted Xiao, Alex Irpan, Mohi Khansari, Dmitry Kalashnikov, and Sergey Levine · 2021
Later among the works it cites.
Calibration-free monocular vision-based robot manipulations with occlusion awareness
Yongle Luo, Kun Dong, Lili Zhao, Zhiyong Sun, Erkang Cheng, Honglin Kan, Chao Zhou, and Bo Song · 2021
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Mart’in-Mart’in · 2021
Later among the works it cites.
A novel accurate positioning method for object pose estimation in robotic manipulation based on vision and tactile sensors
Dan Zhao, Fuchun Sun, Zongtao Wang, and Quan Zhou · 2021
Later among the works it cites.