Fetching the paper…
Reading the bibliography…
A fairly reliable trend in deep reinforcement learning is that the performance scales with the number of parameters, provided a complimentary scaling in amount of training data.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Kernel embeddings of conditional distributions: A unified kernel framework for nonparametric inference in graphical models
L. Song, K. Fukumizu, and A. Gretton · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. Van Hasselt, and D. Silver · 2016
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Earlier work this paper cites.
A. Argenson and G. Dulac-Arnold · 2020
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Earlier work this paper cites.
Generative temporal difference learning for infinite-horizon prediction
M. Janner, I. Mordatch, and S. Levine · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
J. Harb, T. Schaul, D. Precup, and P.-L. Bacon · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
C-learning: Learning to achieve goals via recursive classification
B. Eysenbach, R. Salakhutdinov, and S. Levine · 2020
Cited alongside, same era.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Cited alongside, same era.
Perceiver io: A general architecture for structured inputs & outputs
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, et al · 2021
Low-rank representation of reinforcement learning policies
B. Mazoure, T. Doan, T. Li, V. Makarenkov, J. Pineau, D. Precup, and G. Rabusseau · 2022
Later among the works it cites.
Multi-game decision transformers
K.-H. Lee, O. Nachum, M. S. Yang, L. Lee, D. Freeman, S. Guadarrama, I. Fischer, W. Xu, E. Jang, H. Michalewski, et al · 2022
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
B. Eysenbach, T. Zhang, R. Salakhutdinov, and S. Levine · 2022
Later among the works it cites.
Contrastive value learning: Implicit models for simple offline rl, 2022
B. Mazoure, B. Eysenbach, O. Nachum, and J. Tompson · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Comparing popular simulation environments in the scope of robotics and reinforcement learning
M. Körber, J. Lange, S. Rediske, S. Steinmann, and R. Glück · 2021
Cited alongside, same era.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, D. Hjelm, P. Bachman, and A. Courville · 2021
Cited alongside, same era.
Delayed feedback in episodic reinforcement learning
B. Howson, C. Pike-Burke, and S. Filippi · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Cited alongside, same era.
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Do as i can, not as i say: Grounding language in robotic affordances
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, et al · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.