Fetching the paper…
Reading the bibliography…
Imitating human demonstrations is a promising approach to endow robots with various manipulation capabilities.
A unified approach for motion and force control of robot manipulators: The operational space formulation
O. Khatib · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
Movement imitation with nonlinear dynamical systems in humanoid robots
A. J. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2006
Earlier work this paper cites.
Robot programming by demonstration
A. Billard, S. Calinon, R. Dillmann, and S. Schaal · 2008
Earlier work this paper cites.
Dynamical system modulation for robot learning via kinesthetic demonstrations
M. Hersch, F. Guenter, S. Calinon, and A. Billard · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning and reproduction of gestures by imitation
S. Calinon, F. D’halluin, E. L. Sauser, D. G. Caldwell, and A. Billard · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Imitation learning of positional and force skills demonstrated via kinesthetic teaching and haptic input
P. Kormushev, S. Calinon, and D. G. Caldwell · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective
B. Akgun, M. Cakmak, J. W. Yoo, and A. L. Thomaz · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Open-source benchmarking for learned reaching motion generation in robotics
A. Lemme, Y. Meirovitch, M. Khansari-Zadeh, T. Flash, A. Billard, and J. J. Steil · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. B. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
R. Loftin, B. Peng, J. MacGlashan, M. L. Littman, M. E. Taylor, J. Huang, and D. L. Roberts · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, K. Goldberg, and P. Abbeel · 2017
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Learning robot objectives from physical human interaction
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, D. Roberts, M. E. Taylor, and M. L. Littman · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei · 2018
Cited alongside, same era.
Multiple interactions made easy (mime): Large scale demonstrations data for imitation
P. Sharma, L. Mohan, L. Pinto, and A. Gupta · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Learning from physical human corrections, one feature at a time
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2018
Cited alongside, same era.
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Y. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Cited alongside, same era.
Surreal: Open-source reinforcement learning framework and robot manipulation benchmark
L. Fan, Y. Zhu, J. Zhu, Z. Liu, O. Zeng, A. Gupta, J. Creus-Costa, S. Savarese, and L. Fei-Fei · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
S. Cabi, S. Gómez Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, et al · 2019
Cited alongside, same era.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Cited alongside, same era.
A. Mandlekar, J. Booher, M. Spero, A. Tung, A. Gupta, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu · 2020
Later among the works it cites.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Later among the works it cites.
Human-in-the-loop imitation learning using remote teleoperation
A. Mandlekar, D. Xu, R. Martín-Martín, Y. Zhu, L. Fei-Fei, and S. Savarese · 2020
Later among the works it cites.
What matters in on-policy reinforcement learning? a large-scale empirical study
M. Andrychowicz, A. Raichuk, P. Stańczyk, M. Orsini, S. Girgin, R. Marinier, L. Hussenot, M. Geist, O. Pietquin, M. Michalski, et al · 2020
Later among the works it cites.
Benchmark for skill learning from demonstration: Impact of user experience, task complexity, and start configuration on performance
M. A. Rana, D. Chen, J. Williams, V. Chu, S. R. Ahmadzadeh, and S. Chernova · 2020
Later among the works it cites.
Softgym: Benchmarking deep reinforcement learning for deformable object manipulation
X. Lin, Y. Wang, J. Olkin, and D. Held · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Y. Zhu, J. Wong, A. Mandlekar, and R. Martín-Martín · 2020
Later among the works it cites.
Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
O. Ahmed, F. Träuble, A. Goyal, A. Neitz, Y. Bengio, B. Schölkopf, M. Wüthrich, and S. Bauer · 2020
Later among the works it cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Later among the works it cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Later among the works it cites.
Visual imitation made easy
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto · 2020
Later among the works it cites.
A framework for efficient robotic manipulation
A. Zhan, P. Zhao, L. Pinto, P. Abbeel, and M. Laskin · 2020
Later among the works it cites.
Benchmarks for deep off-policy evaluation
J. Fu, M. Norouzi, O. Nachum, G. Tucker, Z. Wang, A. Novikov, M. Yang, M. R. Zhang, Y. Chen, A. Kumar, et al · 2021
Closest in time.
C. Wang, R. Wang, D. Xu, A. Mandlekar, L. Fei-Fei, and S. Savarese · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Closest in time.
Representation matters: Offline pretraining for sequential decision making
M. Yang and O. Nachum · 2021
Closest in time.
Provable representation learning for imitation with contrastive fourier features
O. Nachum and M. Yang · 2021
Closest in time.
Hyperparameter selection for imitation learning
L. Hussenot, M. Andrychowicz, D. Vincent, R. Dadashi, A. Raichuk, L. Stafiniak, S. Girgin, R. Marinier, N. Momchev, S. Ramos, et al · 2021
Closest in time.