Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A. (2018) · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
Original
Pohlen, T., Piot, B., Hester, T., Azar, M. G., Horgan, D., Budden, D., Barth-Maron, G., Van Hasselt, H., Quan, J., Večerík, M., et al. (2018) · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Deepmind control suite
Original
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A. (2018) · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D. (2018) · 2018
Later among the works it cites.
TF-Agents: A library for reinforcement learning in tensorflow
Guadarrama, S., Korattikara, A., Ramirez, O., Castro, P., Holly, E., Fishman, S., Wang, K., Gonina, E., Wu, N., Harris, C., Vanhoucke, V., and Brevdo, E. (2018) · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N. (2019) · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R. (2019) · 2019
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2019) · 2019
Later among the works it cites.
dm_env: A python interface for reinforcement learning environments
Muldal, A., Doron, Y., Aslanides, J., Harley, T., Ward, T., and Liu, S. (2019) · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W., Mathieu, M., Dudzik, A., Chung, J., Choi, D., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J., Jaderberg, M., and Silver, D. (2019) · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M. (2020) · 2020
Closest in time.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., et al. (2020) · 2020
Closest in time.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. (2020) · 2020
Closest in time.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M. (2020) · 2020
Closest in time.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T., Gómez, S., Zolna, K., Agarwal, R., Merel, J. S., Mankowitz, D. J., Paduraru, C., et al. (2020) · 2020
Closest in time.
Imitation learning as f-divergence minimization
Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S. (2020) · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Closest in time.
Learning dexterous in-hand manipulation
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J. W., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W. (2020) · 2020
Closest in time.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Reddy, S., Dragan, A. D., and Levine, S. (2020) · 2020
Closest in time.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Song, Y., Wang, J., Lukasiewicz, T., Xu, Z., Xu, M., Ding, Z., and Wu, L. (2020) · 2020
Closest in time.
Munchausen reinforcement learning
Vieillard, N., Pietquin, O., and Geist, M. (2020) · 2020
Closest in time.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al. (2020) · 2020
Closest in time.
Fiber: A platform for efficient development and distributed training for reinforcement learning and population-based methods
Zhi, J., Wang, R., Clune, J., and Stanley, K. O. (2020) · 2020
Closest in time.
Offline rl without off-policy evaluation
Brandfonbrener, D., Whitney, W., Ranganath, R., and Bruna, J. (2021) · 2021
Closest in time.
Reverb: a framework for experience replay
Original
Cassirer, A., Barth-Maron, G., Brevdo, E., Ramos, S., Boyd, T., Sottiaux, T., and Kroiss, M. (2021) · 2021
Closest in time.
Primal wasserstein imitation learning
Dadashi, R., Hussenot, L., Geist, M., and Pietquin, O. (2021) · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S. (2021) · 2021
Closest in time.
Regularized behavior value estimation
Original
Gulcehre, C., Colmenarejo, S. G., Wang, Z., Sygnowski, J., Paine, T., Zolna, K., Chen, Y., Hoffman, M., Pascanu, R., and de Freitas, N. (2021) · 2021
Closest in time.
What matters for adversarial imitation learning?
Orsini, M., Raichuk, A., Hussenot, L., Vincent, D., Dadashi, R., Girgin, S., Geist, M., Bachem, O., Pietquin, O., and Andrychowicz, M. (2021) · 2021
Closest in time.
Rlds: an ecosystem to generate, share and use datasets in reinforcement learning
Ramos, S., Girgin, S., Hussenot, L., Vincent, D., Yakubovich, H., Toyama, D., Gergely, A., Stanczyk, P., Marinier, R., Harmsen, J., Pietquin, O., and Momchev, N. (2021) · 2021
Closest in time.
Launchpad: a programming model for distributed machine learning research
Original
Yang, F., Barth-Maron, G., Stańczyk, P., Hoffman, M., Liu, S., Kroiss, M., Pope, A., and Rrustemi, A. (2021) · 2021
Closest in time.
Continuous control with action quantization from demonstrations
Dadashi, R., Hussenot, L., Vincent, D., Girgin, S., Raichuk, A., Geist, M., and Pietquin, O. (2022) · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al. (2022) · 2022
Closest in time.