Fetching the paper…
Reading the bibliography…
We find that across a wide range of robot policy learning scenarios, treating supervised policy learning with an implicit model generally performs better, on average, than commonly used explicit models.
Approximation of continuous and discontinuous functions by generalized sampling series
P. Butzer, S. Ries, and R. Stens · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Robot learning from demonstration
C. G. Atkeson and S. Schaal · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Random projection in dimensionality reduction: applications to image and text data
E. Bingham and H. Mannila · 2001
Earlier work this paper cites.
Neural-network approximation of piecewise continuous functions: application to friction compensation
R. R. Selmic and F. L. Lewis · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
P.-T. De Boer, D. P. Kroese, S. Mannor, and R. Y. Rubinstein · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. Huang · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Approximation of the discontinuities of a function by its classical orthogonal polynomial fourier coefficients
G. Kvernadze · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Accurate reconstruction of discontinuous functions using the singular pade-chebyshev method
A. L. Tampos, J. E. C. Lope, and J. S. Hesthaven · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Model-free imitation learning with policy optimization
J. Ho, J. Gupta, and S. Ermon · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2016
Earlier work this paper cites.
E. Stella, C. Ladera, and G. Donoso · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Optnet: Differentiable optimization as a layer in neural networks
B. Amos and J. Z. Kolter · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Cited alongside, same era.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Your classifier is secretly an energy based model and you should treat it like one
W. Grathwohl, K.-C. Wang, J.-H. Jacobsen, D. Duvenaud, M. Norouzi, and K. Swersky · 2019
Later among the works it cites.
Tossingbot: Learning to throw arbitrary objects with residual physics
A. Zeng, S. Song, J. Lee, A. Rodriguez, and T. Funkhouser · 2020
Later among the works it cites.
Model-based planning with energy-based models
Y. Du, T. Lin, and I. Mordatch · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
How neural networks extrapolate: From feedforward to graph neural networks
K. Xu, M. Zhang, J. Li, S. S. Du, K.-i. Kawarabayashi, and S. Jegelka · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An intriguing failing of convolutional neural networks and the coordconv solution
R. Liu, J. Lehman, P. Molino, F. Petroski Such, E. Frank, A. Sergeev, and J. Yosinski · 2018
Cited alongside, same era.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Cited alongside, same era.
Concept learning with energy-based models
I. Mordatch · 2018
Cited alongside, same era.
Sparsemap: Differentiable sparse structured inference
V. Niculae, A. Martins, M. Blondel, and C. Cardie · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters · 2018
Cited alongside, same era.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne · 2018
Cited alongside, same era.
Later among the works it cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
The surprising effectiveness of linear models for visual foresight in object pile manipulation
H. Suh and R. Tedrake · 2020
Later among the works it cites.
Residual energy-based models for text generation
Y. Deng, A. Bakhtin, M. Ott, A. Szlam, and M. Ranzato · 2020
Later among the works it cites.
Contactnets: Learning discontinuous contact dynamics with smooth, implicit representations
S. Pfrommer, M. Halm, and M. Posa · 2020
Later among the works it cites.
Energy-based imitation learning
M. Liu, T. He, M. Xu, and W. Zhang · 2020
Later among the works it cites.
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. L. Paine, S. G. Colmenarejo, K. Zolna, R. Agarwal, J. Merel, D. Mankowitz, C. Paduraru, et al · 2020
Later among the works it cites.
On the sample complexity of stability constrained imitation learning
S. Tu, A. Robey, T. Zhang, and N. Matni · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
I. Kostrikov, J. Tompson, R. Fergus, and O. Nachum · 2021
Closest in time.
Provable representation learning for imitation with contrastive fourier features
O. Nachum and M. Yang · 2021
Closest in time.
How to train your energy-based models
Y. Song and D. P. Kingma · 2021
Closest in time.
Gradient penalty from a maximum margin perspective
A. Jolicoeur-Martineau and I. Mitliagkas · 2021
Closest in time.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning
S. Sinha and A. Garg · 2021
Closest in time.
Semi-algebraic approximation using christoffel–darboux kernel
S. Marx, E. Pauwels, T. Weisser, D. Henrion, and J. B. Lasserre · 2021
Closest in time.
Improved contrastive divergence training of energy based models
Y. Du, S. Li, B. J. Tenenbaum, and I. Mordatch · 2021
Closest in time.
Amp: Adversarial motion priors for stylized physics-based character control
X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa · 2021
Closest in time.