Fetching the paper…
Reading the bibliography…
Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements.
Bias-variance error bounds for temporal difference updates
Kearns, M. and Singh, S · 2000
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
François-Lavet, V., Fonteneau, R., and Ernst, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Earlier work this paper cites.
Improving regression performance with distributional losses
Imani, E. and White, M · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Better exploration with optimistic actor critic
Ciosek, K., Vuong, Q., Loftin, R., and Hofmann, K · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Cited alongside, same era.
Expected policy gradients for reinforcement learning
Ciosek, K. and Whiteson, S · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Cited alongside, same era.
Cross q q : Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Bhatt, A., Palenicek, D., Belousov, B., Argus, M., Amiranashvili, A., Brox, T., and Peters, J · 2023
Later among the works it cites.
Learning pessimism for reinforcement learning
Cetin, E. and Celiktutan, O · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
What matters in on-policy reinforcement learning? a large-scale empirical study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al · 2021
Cited alongside, same era.
Towards deeper deep reinforcement learning with spectral normalization
Bjorck, N., Gomes, C. P., and Weinberger, K. Q · 2021
Cited alongside, same era.
Dropout q-functions for doubly efficient reinforcement learning
Hiraoka, T., Imagawa, T., Hashimoto, T., Onishi, T., and Tsuruoka, Y · 2021
Cited alongside, same era.
Offline q-learning on diverse multi-task data both scales and generalizes
Kumar, A., Agarwal, R., Geng, X., Tucker, G., and Levine, S · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
Later among the works it cites.
A generalist dynamics model for control
Schubert, I., Zhang, J., Bruce, J., Bechtle, S., Parisotto, E., Riedmiller, M., Springenberg, J. T., Byravan, A., Hasenclever, L., and Heess, N · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
Investigating multi-task pretraining and generalization in reinforcement learning
Taiga, A. A., Agarwal, R., Farebrother, J., Courville, A., and Bellemare, M. G · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al · 2023
Later among the works it cites.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., et al · 2024
Closest in time.
Dissecting deep rl with high update ratios: Combatting value overestimation and divergence
Hussing, M., Voelcker, C., Gilitschenski, I., Farahmand, A.-m., and Eaton, E · 2024
Closest in time.
Slow and steady wins the race: Maintaining plasticity with hare and tortoise networks
Lee, H., Cho, H., Kim, H., Kim, D., Min, D., Choo, J., and Lyle, C · 2024
Closest in time.
Disentangling the causes of plasticity loss in neural networks
Lyle, C., Zheng, Z., Khetarpal, K., van Hasselt, H., Pascanu, R., Martens, J., and Dabney, W · 2024
Closest in time.
On the theory of risk-aware agents: Bridging actor-critic and economics
Nauman, M. and Cygan, M · 2024
Closest in time.
Nauman, M., Bortkiewicz, M., Miłoś, P., Trzcinski, T., Ostaszewski, M., and Cygan, M · 2024
Closest in time.
Small batch deep reinforcement learning
Obando Ceron, J., Bellemare, M., and Castro, P. S · 2024
Closest in time.
Mixtures of experts unlock parameter scaling for deep rl
Obando-Ceron, J., Sokar, G., Willi, T., Lyle, C., Farebrother, J., Foerster, J., Dziugaite, G. K., Precup, D., and Castro, P. S · 2024
Closest in time.
Efficientzero v2: Mastering discrete and continuous control with limited data
Wang, S., Liu, S., Ye, W., You, J., and Gao, Y · 2024
Closest in time.