Fetching the paper…
Reading the bibliography…
Sample efficiency remains a crucial challenge in applying Reinforcement Learning (RL) to real-world tasks.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Rubinstein, R. Y · 1997
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Karnin, Z., Koren, T., and Somekh, O · 2013
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Earlier work this paper cites.
Multi-step reinforcement learning: A unifying algorithm
De Asis, K., Hernandez-Garcia, J., Holland, G., and Sutton, R · 2018
Earlier work this paper cites.
Model-based value expansion for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Cited alongside, same era.
Learning agile and dynamic motor skills for legged robots
Hwangbo, J., Lee, J., Dosovitskiy, A., Bellicoso, D., Tsounis, V., Koltun, V., and Hutter, M · 2019
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Cited alongside, same era.
Stochastic beams and where to find them: The gumbel-top-k trick for sampling sequences without replacement
Kool, W., Van Hoof, H., and Welling, M · 2019
Cited alongside, same era.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Online and offline reinforcement learning by planning with a learned model
Schrittwieser, J., Hubert, T., Mandhane, A., Barekatain, M., Antonoglou, I., and Silver, D · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Mastering atari games with limited data
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y · 2021
Later among the works it cites.
Improving deep neural networks using softplus units
Zheng, H., Yang, Z., Liu, W., Liang, J., and Li, Y · 2021
Later among the works it cites.
A system for general in-hand object re-orientation
Chen, T., Xu, J., and Agrawal, P · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Planning in stochastic environments with a learned model
Antonoglou, I., Schrittwieser, J., Ozair, S., Hubert, T. K., and Silver, D · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Cited alongside, same era.
Policy improvement by planning with gumbel
Danihelka, I., Guez, A., Schrittwieser, J., and Silver, D · 2021
Cited alongside, same era.
Hansen, N., Wang, X., and Su, H · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning, 2022
Wu, P., Escontrela, A., Hafner, D., Goldberg, K., and Abbeel, P · 2022
Later among the works it cites.
Efficient learning for alphazero via path consistency
Zhao, D., Tu, S., and Xu, L · 2022
Later among the works it cites.
Visual dexterity: In-hand reorientation of novel and complex object shapes
Chen, T., Tippur, M., Wu, S., Kumar, V., Adelson, E., and Agrawal, P · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control, 2023
Hansen, N., Su, H., and Wang, X · 2023
Later among the works it cites.
Dexpbt: Scaling up dexterous manipulation for hand-arm systems with population based training
Petrenko, A., Allshire, A., State, G., Handa, A., and Makoviychuk, V · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.