Fetching the paper…
Reading the bibliography…
Reinforcement learning research obtained significant success and attention with the utilization of deep neural networks to solve problems in high dimensional state or action spaces.
Dynamic programming
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Functional approximation and dynamic programming
Bellman, R. and Dreyfus, S. (1959) · 1959
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. (1984) · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal difference
Sutton, R. (1988) · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. (1989) · 1989
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, L.-J. (1993) · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A. (1993) · 1993
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. A. and Moore, A. W. (1994) · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
III, L. C. B. (1995) · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J. (2000) · 2000
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. (2003) · 2003
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016) · 2003
Earlier work this paper cites.
Cross-domain transfer for reinforcement learning
Taylor, M. E. and Stone, P. (2007) · 2007
Earlier work this paper cites.
Double q-learning
van Hasselt, H. (2010) · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2014) · 2014
Earlier work this paper cites.
Explaning and harnessing adversarial examples
Goodfellow, I., Shelens, J., and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, a. G., Graves, A., Riedmiller, M., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hasselt, H. v., Guez, A., and Silver, D. (2016) · 2016
Earlier work this paper cites.
VIME: variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Roy, B. V. (2016a) · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N. (2017) · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Silver, D., and van Hasselt, H. (2017) · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S., Papernot, N., Goodfellow, Ian an Duan, Y., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Delving into adversarial attacks on deep policies
Kos, J. and Song, D. (2017) · 2017
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Lin, Y.-C., Zhang-Wei, H., Liao, Y.-H., Shih, M.-L., Liu, i.-Y., and Sun, M. (2017) · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A. (2017) · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T. P., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017) · 2017
Munchausen reinforcement learning
Vieillard, N., Pietquin, O., and Geist, M. (2020b) · 2020
Later among the works it cites.
Improving generalization in reinforcement learning with mixture regularization
Wang, K., Kang, B., Shao, J., and Feng, J. (2020) · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H. P., Hessel, M., Oh, J., Singh, S., and Silver, D. (2020) · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. G. (2021b) · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
Garg, D., Chakraborty, S., Cundy, C., Song, J., and Ermon, S. (2021) · 2021
Later among the works it cites.
Investigating vulnerabilities of deep neural policies
Korkmaz, E. (2021b) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S. (2018) · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D. (2018) · 2018
Cited alongside, same era.
Transfer of value functions via variational methods
Tirinzoni, A., Rodríguez-Sánchez, R., and Restelli, M. (2018) · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J. (2019) · 2019
Cited alongside, same era.
Transfer learning for related reinforcement learning tasks via image-to-image translation
Gamrian, S. and Goldberg, Y. (2019) · 2019
Cited alongside, same era.
A meta-mdp approach to exploration for lifelong reinforcement learning
Garcia, F. M. and Thomas, P. S. (2019) · 2019
Cited alongside, same era.
Later among the works it cites.
Metrics and continuity in reinforcement learning
Lan, C. L., Bellemare, M. G., and Castro, P. S. (2021) · 2021
Later among the works it cites.
Lipschitz lifelong reinforcement learning
Lecarpentier, E., Abel, D., Asadi, K., Jinnai, Y., Rachelson, E., and Littman, M. L. (2021) · 2021
Later among the works it cites.
SUNRISE: A simple unified framework for ensemble learning in deep reinforcement learning
Lee, K., Laskin, M., Srinivas, A., and Abbeel, P. (2021) · 2021
Later among the works it cites.
Regularization matters in policy optimization - an empirical study on continuous control
Liu, Z., Li, X., and Darrell, T. (2021) · 2021
Later among the works it cites.
When is generalizable reinforcement learning tractable?
Malik, D., Li, Y., and Ravikumar, P. (2021) · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Raileanu, R. and Fergus, R. (2021) · 2021
Later among the works it cites.
Discovery of options via meta-learned subgoals
Veeriah, V., Zahavy, T., Hessel, M., Xu, Z., Oh, J., Kemaev, I., van Hasselt, H., Silver, D., and Singh, S. (2021) · 2021
Later among the works it cites.
Continual world: A robotic benchmark for continual reinforcement learning
Wolczyk, M., Zajac, M., Pascanu, R., Kucinski, L., and Milos, P. (2021) · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R. (2021) · 2021
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., Ruiz, F. J. R., Schrittwieser, J., Swirszcz, G., Silver, D., Hassabis, D., and Kohli, P. (2022) · 2022
Later among the works it cites.
Introducing symmetries to black box meta reinforcement learning
Kirsch, L., Flennerhag, S., van Hasselt, H., Friesen, A. L., Oh, J., and Chen, Y. (2022) · 2022
Later among the works it cites.
Deep reinforcement learning policies learn shared adversarial features across mdps
Korkmaz, E. (2022) · 2022
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Google Gemini (2023) · 2023
Later among the works it cites.
Human-level atari 200x faster
Kapturowski, S., Campos, V., Jiang, R., Rakicevic, N., van Hasselt, H., Blundell, C., and Badia, A. P. (2023) · 2023
Later among the works it cites.
Adversarial robust deep reinforcement learning requires redefining robustness
Korkmaz, E. (2023) · 2023
Later among the works it cites.
Detecting adversarial directions in deep reinforcement learning to make robust decisions
Korkmaz, E. and Brown-Cohen, J. (2023) · 2023
Later among the works it cites.
Faster sorting algorithms discovered using deep reinforcement learning
Mankowitz, D. J., Michi, A., Zhernov, A., Gelmi, M., Selvi, M., Paduraru, C., Leurent, E., Iqbal, S., Lespiau, J., Ahern, A., Köppe, T., Millikin, K., Gaffney, S., Elster, S., Broshear, J., Gamble, C., Milan, K., Tung, R., Hwang, M., Cemgil, T., Barekatain, M., Li, Y., Mandhane, A., Hubert, T., Schrittwieser, J., Hassabis, D., Kohli, P., Riedmiller, M. A., Vinyals, O., and Silver, D. (2023) · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI (2023) · 2023
Later among the works it cites.
Random latent exploration for deep reinforcement learning
Mahankali, S., Hong, Z., Sekhari, A., Rakhlin, A., and Agrawal, P. (2024) · 2024
Closest in time.
Phasic policy gradient
Cobbe, K., Hilton, J., Klimov, O., and Schulman, J. (2021) · 2027
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Closest in time.