Fetching the paper…
Reading the bibliography…
We show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Learning by playing solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Human compatible: Artificial intelligence and the problem of control
Russell, S · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Cited alongside, same era.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al · 2020
Cited alongside, same era.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Chebotar, Y., Hausman, K., Lu, Y., Xiao, T., Kalashnikov, D., Varley, J., Irpan, A., Eysenbach, B., Julian, R., Finn, C., et al · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Later among the works it cites.
Scaling laws for a multi-agent reinforcement learning model
Neumann, O. and Gros, C · 2022
Later among the works it cites.
A generalist agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-maron, G., Giménez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Later among the works it cites.
Robocat: A self-improving foundation agent for robotic manipulation
Bousmalis, K., Vezzani, G., Rao, D., Devin, C., Lee, A. X., Bauza, M., Davchev, T., Zhou, Y., Gupta, A., Raju, A., et al · 2023
Later among the works it cites.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
Chebotar, Y., Vuong, Q., Hausman, K., Xia, F., Lu, Y., Irpan, A., Kumar, A., Yu, T., Herzog, A., Pertsch, K., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Cited alongside, same era.
Beyond pick-and-place: Tackling robotic stacking of diverse shapes
Lee, A. X., Devin, C., Zhou, Y., Lampe, T., Bousmalis, K., Springenberg, J. T., Byravan, A., Abdolmaleki, A., Gileadi, N., Khosid, D., Fantacci, C., Chen, J. E., Raju, A., Jeong, R., Neunert, M., Laurens, A., Saliceti, S., Casarini, F., Riedmiller, M., Hadsell, R., and Nori, F · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Cited alongside, same era.
On multi-objective policy optimization as a tool for reinforcement learning: Case studies in offline rl and finetuning, 2022
Abdolmaleki, A., Huang, S., Vezzani, G., Shahriari, B., Springenberg, J. T., Mishra, S., Tirumala, D., Byravan, A., Bousmalis, K., György, A., et al · 2022
Cited alongside, same era.
Vision-language models as success detectors
Du, Y., Konyushkova, K., Denil, M., Raju, A., Landon, J., Hill, F., de Freitas, N., and Cabi, S · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X · 2023
Later among the works it cites.
Scaling laws for single-agent reinforcement learning
Hilton, J., Tang, J., and Schulman, J · 2023
Later among the works it cites.
Scaling up gans for text-to-image synthesis
Kang, M., Zhu, J.-Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T · 2023
Later among the works it cites.
Mastering stacking of diverse shapes with large-scale iterative reinforcement learning on real robots, 2023
Lampe, T., Abdolmaleki, A., Bechtle, S., Huang, S. H., Springenberg, J. T., Bloesch, M., Groth, O., Hafner, R., Hertweck, T., Neunert, M., Wulfmeier, M., Zhang, J., Nori, F., Heess, N., and Riedmiller, M · 2023
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Xu, C., Luo, J., Kreiman, T., Tan, Y., Sadigh, D., Finn, C., and Levine, S · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Open X-Embodiment Collaboration · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Gnfactor: Multi-task real robot learning with generalizable neural feature fields
Ze, Y., Yan, G., Wu, Y.-H., Macaluso, A., Ge, Y., Ye, J., Hansen, N., Li, L. E., and Wang, X · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al · 2023
Later among the works it cites.