Fetching the paper…
Reading the bibliography…
A key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization.
Gradient theory of optimal flight paths
Kelley, H. J · 1960
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Schmidhuber, Jürgen · 1990
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Image-based simultaneous control of robot and target object motions by direct-image-interpretation method
Deguchi, Koichiro and Takahashi, Isao · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, Andrew Y, Harada, Daishi, and Russell, Stuart · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, Andrew Y and Russell, Stuart · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, Pieter and Ng, Andrew Y · 2004
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
Lange, Sascha, Riedmiller, Martin, and Voigtlander, Arne · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Yuval, Erez, Tom, and Todorov, Emanuel · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, Diederik P and Welling, Max · 2013
Earlier work this paper cites.
Learning state representations with robotic priors
Jonschkowski, Rico and Brock, Oliver · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, Alec, Metz, Luke, and Chintala, Soumith · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, Manuel, Springenberg, Jost, Boedecker, Joschka, and Riedmiller, Martin · 2015
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S, Davis, Andy, Dean, Jeffrey, Devin, Matthieu, et al · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, Pulkit, Nair, Ashvin V, Abbeel, Pieter, Malik, Jitendra, and Levine, Sergey · 2016
Earlier work this paper cites.
Ba, Jimmy Lei, Kiros, Jamie Ryan, and Hinton, Geoffrey E · 2016
Earlier work this paper cites.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Finn, Chelsea, Tan, Xin Yu, Duan, Yan, Darrell, Trevor, Levine, Sergey, and Abbeel, Pieter · 2016
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, Ian, Bengio, Yoshua, and Courville, Aaron · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, Jonathan and Ermon, Stefano · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2016
Cited alongside, same era.
The curious robot: Learning visual representations via physical interactions
Pinto, Lerrel, Gandhi, Dhiraj, Han, Yuanfeng, Park, Yong-Lae, and Gupta, Abhinav · 2016
Cited alongside, same era.
Cad2rl: Real single-image flight without a single real image
Sadeghi, Fereshteh and Levine, Sergey · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Darla: Improving zero-shot transfer in reinforcement learning
Higgins, Irina, Pal, Arka, Rusu, Andrei A, Matthey, Loic, Burgess, Christopher P, Pritzel, Alexander, Botvinick, Matthew, Blundell, Charles, and Lerchner, Alexander · 2017
Later among the works it cites.
Pves: Position-velocity encoders for unsupervised learning of structured state representations
Jonschkowski, Rico, Hafner, Roland, Scholz, Jonathan, and Riedmiller, Martin · 2017
Later among the works it cites.
Inferring the latent structure of human decision-making from raw visual inputs
Li, Yunzhu, Song, Jiaming, and Ermon, Stefano · 2017
Later among the works it cites.
Combining self-supervised learning and imitation for vision-based rope manipulation
Nair, Ashvin, Chen, Dian, Agrawal, Pulkit, Isola, Phillip, Abbeel, Pieter, Malik, Jitendra, and Levine, Sergey · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Salimans, Tim and Kingma, Diederik P · 2016
Cited alongside, same era.
Optimizing Expectations: From Deep Reinforcement Learning to Stochastic Computation Graphs
Schulman, John · 2016
Cited alongside, same era.
Unsupervised perceptual rewards for imitation learning
Sermanet, Pierre, Xu, Kelvin, and Levine, Sergey · 2016
Cited alongside, same era.
The predictron: End-to-end learning and planning
Silver, David, van Hasselt, Hado, Hessel, Matteo, Schaul, Tom, Guez, Arthur, Harley, Tim, Dulac-Arnold, Gabriel, Reichert, David, Rabinowitz, Neil, Barreto, Andre, et al · 2016
Cited alongside, same era.
Value iteration networks
Tamar, Aviv, Wu, Yi, Thomas, Garrett, Levine, Sergey, and Abbeel, Pieter · 2016
Cited alongside, same era.
Incorporating human domain knowledge into large scale cost function learning
Wulfmeier, Markus, Rao, Dushyant, and Posner, Ingmar · 2016
Cited alongside, same era.
Optnet: Differentiable optimization as a layer in neural networks
Amos, Brandon and Kolter, J Zico · 2017
Cited alongside, same era.
Value prediction network
Oh, Junhyuk, Singh, Satinder, and Lee, Honglak · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, Deepak, Agrawal, Pulkit, Efros, Alexei A., and Darrell, Trevor · 2017
Later among the works it cites.
Swish: a self-gated activation function
Ramachandran, Prajit, Zoph, Barret, and Le, Quoc V · 2017
Later among the works it cites.
Openai mujoco-py, 2017
Schneider, Jonas, Welinder, Peter, Ray, Alex, Ho, Jonathan, and Zaremba, Wojceich · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from multi-view observation
Sermanet, Pierre, Lynch, Corey, Hsu, Jasmine, and Levine, Sergey · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, Sainbayar, Kostrikov, Ilya, Szlam, Arthur, and Fergus, Rob · 2017
Later among the works it cites.
Learning from the hindsight plan—episodic mpc improvement
Tamar, Aviv, Thomas, Garrett, Zhang, Tianhao, Levine, Sergey, and Abbeel, Pieter · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, Josh, Fong, Rachel, Ray, Alex, Schneider, Jonas, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, Théophane, Racanière, Sébastien, Reichert, David P, Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adria Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, et al · 2017
Later among the works it cites.
Integrating state representation learning into deep reinforcement learning
de Bruin, T., Kober, J., Tuyls, K., and Babuska, R · 2018
Closest in time.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Finn, Chelsea and Levine, Sergey · 2018
Closest in time.
META LEARNING SHARED HIERARCHIES
Frans, Kevin, Ho, Jonathan, Chen, Xi, Abbeel, Pieter, and Schulman, John · 2018
Closest in time.
Learning to search with MCTSnets, 2018
Guez, Arthur, Weber, Theophane, Antonoglou, Ioannis, Simonyan, Karen, Vinyals, Oriol, Wierstra, Daan, Munos, Remi, and Silver, David · 2018
Closest in time.
Zero-shot visual imitation
Pathak*, Deepak, Mahmoudieh*, Parsa, Luo*, Michael, Agrawal*, Pulkit, Chen, Dian, Shentu, Fred, Shelhamer, Evan, Malik, Jitendra, Efros, Alexei A., and Darrell, Trevor · 2018
Closest in time.
Mpc-inspired neural network policies for sequential decision making
Pereira, Marcus, Fan, David D, An, Gabriel Nakajima, and Theodorou, Evangelos · 2018
Closest in time.