Fetching the paper…
Reading the bibliography…
We propose Scheduled Auxiliary Control (SAC-X), a new learning paradigm in the context of Reinforcement Learning (RL).
Improving generalization for temporal difference learning: The successor representation
Dayan, Peter · 1993
Earlier work this paper cites.
Multitask learning
Caruana, Rich · 1997
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
Dietterich, Thomas G · 1998
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Randløv, Jette and Alstrøm, Preben · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, Andrew Y., Harada, Daishi, and Russell, Stuart J · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, Andrew Y. and Russell, Stuart · 2000
Earlier work this paper cites.
Intrinsically Motivated reinforcement learning
Chentanez, Nuttapong, Barto, Andrew G., and Singh, Satinder P · 2005
Earlier work this paper cites.
Transfer of samples in batch reinforcement learning
Lazaric, Alessandro, Restelli, Marcello, and Bonarini, Andrea · 2008
Earlier work this paper cites.
Transfer in variable-reward hierarchical reinforcement learning
Mehta, Neville, Natarajan, Sriraam, Tadepalli, Prasad, and Fern, Alan · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, Brian D., Maas, Andrew L., Bagnell, J. Andrew, and Dey, Anind K · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Where do rewards come from?
Singh, S., Lewis, R. L., and Barto, A. G · 2009
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, J. and Peters, J · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stéphane, Gordon, Geoffrey J., and Bagnell, Drew · 2011
Earlier work this paper cites.
Horde : A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, Richard S., Modayil, Joseph, Delp, Michael, Degris, Thomas, Pilarski, Patrick M., White, Adam, and Precup, Doina · 2011
Earlier work this paper cites.
Hierarchical relative entropy policy search
Daniel, C., Neumann, G., and Peters, J · 2012
Earlier work this paper cites.
Learning skills from play: Artificial curiosity on a katana robot arm
Ngo, Hung Quoc, Luciw, Matthew D., Förster, Alexander, and Schmidhuber, Jürgen · 2012
Earlier work this paper cites.
Better Generalization with Forecasts
Schaul, Tom and Ring, Mark B · 2013
Cited alongside, same era.
Schmidhuber, Jürgen · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, Diederik P and Welling, Max · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, Danilo Jimenez, Mohamed, Shakir, and Wierstra, Daan · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, Djork-Arné, Unterthiner, Thomas, and Hochreiter, Sepp · 2015
Cited alongside, same era.
The Intentional Unintentional Agent : Learning to solve many continuous control tasks simultaneously
Cabi, Serkan, Colmenarejo, Sergio Gomez, Hoffman, Matthew W., Denil, Misha, Wang, Ziyu, and de Freitas, Nando · 2017
Later among the works it cites.
Feature control as intrinsic motivation for hierarchical reinforcement learning
Dilokthanakul, Nat, Kaplanis, Christos, Pawlowski, Nick, and Shanahan, Murray · 2017
Later among the works it cites.
Learning to act by predicting the future
Dosovitskiy, Alexey and Koltun, Vladlen · 2017
Later among the works it cites.
One-shot imitation learning
Duan, Yan, Andrychowicz, Marcin, Stadie, Bradly C., Ho, Jonathan, Schneider, Jonas, Sutskever, Ilya, Abbeel, Pieter, and Zaremba, Wojciech · 2017
Later among the works it cites.
Intrinsically motivated goal exploration processes with automatic curriculum learning
Forestier, Sébastien, Mollard, Yoan, and Oudeyer, Pierre-Yves · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heess, Nicolas, Wayne, Gregory, Silver, David, Lillicrap, Tim, Erez, Tom, and Tassa, Yuval · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P and Ba, Jimmy · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Universal Value function approximators
Schaul, Tom, Horgan, Daniel, Gregor, Karol, and Silver, David · 2015
Cited alongside, same era.
Ba, Lei Jimmy, Kiros, Ryan, and Hinton, Geoffrey E · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Cited alongside, same era.
Learning to Navigate in complex environments
Mirowski, Piotr, Pascanu, Razvan, Viola, Fabio, Soyer, Hubert, Ballard, Andrew J., Banino, Andrea, Denil, Misha, Goroshin, Ross, Sifre, Laurent, Kavukcuoglu, Koray, Kumaran, Dharshan, and Hadsell, Raia · 2016
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, Shixiang, Holly, Ethan, Lillicrap, Timothy, and Levine, Sergey · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Heess, Nicolas, TB, Dhruva, Sriram, Srinivasan, Lemmon, Jay, Merel, Josh, Wayne, Greg, Tassa, Yuval, Erez, Tom, Wang, Ziyu, Eslami, S. M. Ali, Riedmiller, Martin A., and Silver, David · 2017
Later among the works it cites.
Unreal : Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, V, Czarnecki, W M, Schaul, T, Leibo, J Z, Silver, D, and Kavukcuoglu, K · 2017
Later among the works it cites.
Playing FPS games with deep reinforcement learning
Lample, Guillaume and Chaplot, Devendra Singh · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, Ashvin, McGrew, Bob, Andrychowicz, Marcin, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Sim-to-real robot learning from pixels with progressive nets
Rusu, Andrei A., Vecerik, Matej, Rothörl, Thomas, Heess, Nicolas, Pascanu, Razvan, and Hadsell, Raia · 2017
Later among the works it cites.
Sim2real view invariant visual servoing by recurrent control
Sadeghi, Fereshteh, Toshev, Alexander, Jang, Eric, and Levine, Sergey · 2017
Later among the works it cites.
Unsupervised perceptual rewards for imitation learning
Sermanet, Pierre, Xu, Kelvin, and Levine, Sergey · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, Joshua, Fong, Rachel, Ray, Alex, Schneider, Jonas, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, Matej, Hester, Todd, Scholz, Jonathan, Wang, Fumin, Pietquin, Olivier, Piot, Bilal, Heess, Nicolas, Rothörl, Thomas, Lampe, Thomas, and Riedmiller, Martin A · 2017
Later among the works it cites.
Divide-and-conquer reinforcement learning
Ghosh, Dibya, Singh, Avi, Rajeswaran, Aravind, Kumar, Vikash, and Levine, Sergey · 2018
Closest in time.
Distributed prioritized experience replay
Horgan, Dan, Quan, John, Budden, David, Barth-Maron, Gabriel, Hessel, Matteo, van Hasselt, Hado, and Silver, David · 2018
Closest in time.