Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (RL) has emerged as a powerful paradigm to solve a range of complex yet specific control tasks.
Nearest neighbor estimates of entropy
Singh, Harshinder, Misra, Neeraj, Hnizdo, Vladimir, Fedorowicz, Adam, and Demchuk, Eugene · 2003
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, Pierre-Yves, Kaplan, Frdric, and Hafner, Verena V · 2007
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, Diederik P and Welling, Max · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael, and Moritz, Philipp · 2015
Earlier work this paper cites.
Beattie, Charles, Leibo, Joel Z, Teplyashin, Denis, Ward, Tom, Wainwright, Marcus, Küttler, Heinrich, Lefrancq, Andrew, Green, Simon, Valdés, Víctor, Sadik, Amir, et al · 2016
Earlier work this paper cites.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter · 2016
Earlier work this paper cites.
Openai baselines
Dhariwal, Prafulla, Hesse, Christopher, Klimov, Oleg, Nichol, Alex, Plappert, Matthias, Radford, Alec, Schulman, John, Sidor, Szymon, Wu, Yuhuai, and Zhokhov, Peter · 2017
Earlier work this paper cites.
Variational intrinsic control
Gregor, Karol, Rezende, Danilo Jimenez, and Wierstra, Daan · 2017
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z., Silver, David, and Kavukcuoglu, Koray · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, Deepak, Agrawal, Pulkit, Efros, Alexei A, and Darrell, Trevor · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, David, Schrittwieser, Julian, Simonyan, Karen, Antonoglou, Ioannis, Huang, Aja, Guez, Arthur, Hubert, Thomas, Baker, Lucas, Lai, Matthew, Bolton, Adrian, et al · 2017
Earlier work this paper cites.
Variational option discovery algorithms
Achiam, Joshua, Edwards, Harrison, Amodei, Dario, and Abbeel, Pieter · 2018
Earlier work this paper cites.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Barreto, Andre, Borsa, Diana, Quan, John, Schaul, Tom, Silver, David, Hessel, Matteo, Mankowitz, Daniel, Zidek, Augustin, and Munos, Remi · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, Tuomas, Zhou, Aurick, Abbeel, Pieter, and Levine, Sergey · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, Matteo, Modayil, Joseph, van Hasselt, Hado, Schaul, Tom, Ostrovski, Georg, Dabney, Will, Horgan, Dan, Piot, Bilal, Azar, Mohammad Gheshlaghi, and Silver, David · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, David, Hubert, Thomas, Schrittwieser, Julian, Antonoglou, Ioannis, Lai, Matthew, Guez, Artfhur, Lanctot, Marc, Sifre, Laurent, Kumaran, Dharshan, Graepel, Thore, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, Richard S and Barto, Andrew G · 2018
Cited alongside, same era.
Tassa, Yuval, Doron, Yotam, Muldal, Alistair, Erez, Tom, Li, Yazhe, Casas, Diego de Las, Budden, David, Abdolmaleki, Abbas, Merel, Josh, Lefrancq, Andrew, et al · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Akkaya, Ilge, Andrychowicz, Marcin, Chociej, Maciek, Litwin, Mateusz, McGrew, Bob, Petron, Arthur, Paino, Alex, Plappert, Matthias, Powell, Glenn, Ribas, Raphael, et al · 2019
Ecoffet, Adrien, Huizinga, Joost, Lehman, Joel, Stanley, Kenneth O., and Clune, Jeff · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, Justin, Kumar, Aviral, Nachum, Ofir, Tucker, George, and Levine, Sergey · 2020
Later among the works it cites.
Adversarial policies: Attacking deep reinforcement learning
Gleave, Adam, Dennis, Michael, Wild, Cody, Kant, Neel, Levine, Sergey, and Russell, Stuart · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, Jean-Bastien, Strub, Florian, Altché, Florent, Tallec, Corentin, Richemond, Pierre H, Buchatskaya, Elena, Doersch, Carl, Pires, Bernardo Avila, Guo, Zhaohan Daniel, Azar, Mohammad Gheshlaghi, et al · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, Caglar, Wang, Ziyu, Novikov, Alexander, Paine, Tom Le, Colmenarejo, Sergio Gomez, Zolna, Konrad, Agarwal, Rishabh, Merel, Josh, Mankowitz, Daniel, Paduraru, Cosmin, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, Christopher, Brockman, Greg, Chan, Brooke, Cheung, Vicki, Dębiak, Przemysław, Dennison, Christy, Farhi, David, Fischer, Quirin, Hashme, Shariq, Hesse, Chris, et al · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, Karl, Klimov, Oleg, Hesse, Chris, Kim, Taehoon, and Schulman, John · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Chang, Ming-Wei, Lee, Kenton, and Toutanova, Kristina · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, Benjamin, Gupta, Abhishek, Ibarz, Julian, and Levine, Sergey · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, Danijar, Lillicrap, Timothy, Fischer, Ian, Villegas, Ruben, Ha, David, Lee, Honglak, and Davidson, James · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lee, Lisa, Eysenbach, Benjamin, Parisotto, Emilio, Xing, Eric P., Levine, Sergey, and Salakhutdinov, Ruslan · 2019
Cited alongside, same era.
Self-supervised exploration via disagreement
Pathak, Deepak, Gandhi, Dhiraj, and Gupta, Abhinav · 2019
Cited alongside, same era.
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, Danijar, Lillicrap, Timothy, Ba, Jimmy, and Norouzi, Mohammad · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, Steven, Dabney, Will, Barreto, André, Warde-Farley, David, de Wiele, Tom Van, and Mnih, Volodymyr · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, Kaiming, Fan, Haoqi, Wu, Yuxin, Xie, Saining, and Girshick, Ross B · 2020
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
Hénaff, Olivier J., Srinivas, Aravind, Fauw, Jeffrey De, Razavi, Ali, Doersch, Carl, Eslami, S. M. Ali, and van den Oord, Aäron · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, Archit, Gu, Shixiang, Levine, Sergey, Kumar, Vikash, and Hausman, Karol · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, Tianhe, Quillen, Deirdre, He, Zhanpeng, Julian, Ryan, Hausman, Karol, Finn, Chelsea, and Levine, Sergey · 2020
Later among the works it cites.
Beyond fine-tuning: Transferring behavior in reinforcement learning
Campos, Víctor, Sprechmann, Pablo, Hansen, Steven Stenberg, Barreto, Andre, Kapturowski, Steven, Vitvitskyi, Alex, Badia, Adria Puigdomenech, and Blundell, Charles · 2021
Closest in time.
Decision transformer: Reinforcement learning via sequence modeling, 2021
Chen, Lili, Lu, Kevin, Rajeswaran, Aravind, Lee, Kimin, Grover, Aditya, Laskin, Michael, Abbeel, Pieter, Srinivas, Aravind, and Mordatch, Igor · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
Janner, Michael, Li, Qiyang, and Levine, Sergey · 2021
Closest in time.
B-pref: Benchmarking preference-based reinforcement learning
Lee, Kimin, Smith, Laura, Dragan, Anca, and Abbeel, Pieter · 2021
Closest in time.
A policy gradient method for task-agnostic exploration
Mutti, Mirco, Pratissoli, Lorenzo, and Restelli, Marcello · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, Max, Anand, Ankesh, Goel, Rishab, Hjelm, R Devon, Courville, Aaron, and Bachman, Philip · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Seo, Younggyo, Chen, Lili, Shin, Jinwoo, Lee, Honglak, Abbeel, Pieter, and Lee, Kimin · 2021
Closest in time.
Unsupervised learning for reinforcement learning, 2021
Srinivas, Aravind and Abbeel, Pieter · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
Stooke, Adam, Lee, Kimin, Abbeel, Pieter, and Laskin, Michael · 2021
Closest in time.