Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has drawn increasing interests in recent years due to its tremendous success in various applications.
Functional analysis
Kôsaku Yosida et al · 1971
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Maxout networks
Ian Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Chaos, fractals, and noise: stochastic aspects of dynamics , volume 97
Andrzej Lasota and Michael C Mackey · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Kristof Van Moffaert and Ann Nowé · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
R l 2 Rl^{2} : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, David Silver, and Hado P van Hasselt · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Andre Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel Mankowitz, Augustin Zidek, and Remi Munos · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado P van Hasselt, and David Silver · 2018
Cited alongside, same era.
Universal successor features approximators
Diana Borsa, Andre Barreto, John Quan, Daniel J. Mankowitz, Hado van Hasselt, Remi Munos, David Silver, and Tom Schaul · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Fast reinforcement learning with generalized policy updates
André Barreto, Shaobo Hou, Diana Borsa, David Silver, and Doina Precup · 2020
Later among the works it cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Marc G Bellemare, Salvatore Candido, Pablo Samuel Castro, Jun Gong, Marlos C Machado, Subhodeep Moitra, Sameera S Ponda, and Ziyu Wang · 2020
Later among the works it cites.
Max-plus linear approximations for deterministic continuous-state markov decision processes
Eloïse Berthier and Francis Bach · 2020
Later among the works it cites.
Minimax confidence interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah · 2019
Cited alongside, same era.
Multi-task batch reinforcement learning with metric learning
Jiachen Li, Quan Vuong, Shuang Liu, Minghua Liu, Kamil Ciosek, Henrik Iskov Christensen, and Hao Su · 2019
Cited alongside, same era.
Lu Lu, Pengzhan Jin, and George Em Karniadakis · 2019
Cited alongside, same era.
Black-box off-policy estimation for infinite-horizon reinforcement learning
Ali Mousavi, Lihong Li, Qiang Liu, and Denny Zhou · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Cited alongside, same era.
Composing value functions in reinforcement learning
Benjamin Van Niekerk, Steven James, Adam Earle, and Benjamin Rosman · 2019
Cited alongside, same era.
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2020
Later among the works it cites.
A boolean task algebra for reinforcement learning
Geraud Nangue Tasse, Steven James, and Benjamin Rosman · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin Yang, and Russ R Salakhutdinov · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Later among the works it cites.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, ziyu wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, Cosmin Paduraru, Sergey Levine, and Thomas Paine · 2021
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2021
Later among the works it cites.