Fetching the paper…
Reading the bibliography…
In this work, we present a scalable reinforcement learning method for training multi-task policies from large offline datasets that can leverage both human demonstrations and autonomously collected data.
Bail: Best-action imitation learning for batch deep reinforcement learning, 2019
X. Chen, Z. Zhou, Z. Wang, C. Wang, Y. Wu, and K. Ross · 1910
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, D. Kumaran, and R. Hadsell · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel · 2017
Earlier work this paper cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Discrete sequential prediction of continuous actions for deep rl
L. Metz, J. Ibarz, N. Jaitly, and J. Davidson · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson, A. X. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, and A. R. Zamir · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, et al · 2019
Earlier work this paper cites.
Scene memory transformer for embodied agents in long-horizon tasks
K. Fang, A. Toshev, L. Fei-Fei, and S. Savarese · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
A. Kumar, X. B. Peng, and S. Levine · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
R. K. Srivastava, P. Shyam, F. Mutz, W. Jaśkowski, and J. Schmidhuber · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Cited alongside, same era.
Motion reasoning for goal-based imitation learning
D.-A. Huang, Y.-W. Chao, C. Paxton, X. Deng, L. Fei-Fei, J. C. Niebles, A. Garg, and D. Fox · 2020
Cited alongside, same era.
Transformers for one-shot visual imitation
S. Dasari and A. Gupta · 2021
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2022
Later among the works it cites.
Metamorph: Learning universal controllers with transformers
A. Gupta, L. Fan, S. Ganguli, and L. Fei-Fei · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Y. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, and M. Riedmiller · 2020
Cited alongside, same era.
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2020
Cited alongside, same era.
Thinking while moving: Deep reinforcement learning with concurrent control
T. Xiao, E. Jang, D. Kalashnikov, S. Levine, J. Ibarz, K. Hausman, and A. Herzog · 2020
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Cited alongside, same era.
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
N. M. M. Shafiullah, Z. J. Cui, A. Altanzaya, and L. Pinto · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Do as I can, not as I say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Later among the works it cites.
Gnm: A general navigation model to drive any robot
D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine · 2022
Later among the works it cites.
Pre-training for robots: Offline rl enables learning new tasks from a handful of trials
A. Kumar, A. Singh, F. Ebert, Y. Yang, C. Finn, and S. Levine · 2022
Later among the works it cites.
Monte carlo augmented actor-critic for sparse reward deep reinforcement learning from suboptimal demonstrations
A. Wilcox, A. Balakrishna, J. Dedieu, W. Benslimane, D. S. Brown, and K. Goldberg · 2022
Later among the works it cites.
Multi-game decision transformers
K.-H. Lee, O. Nachum, M. Yang, L. Lee, D. Freeman, W. Xu, S. Guadarrama, I. Fischer, E. Jang, H. Michalewski, et al · 2022
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
H. Furuta, Y. Matsuo, and S. S. Gu · 2022
Later among the works it cites.
GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Y. Jang, J. Lee, and K.-E. Kim · 2022
Later among the works it cites.
Offline pre-trained multi-agent decision transformer, 2022
L. Meng, M. Wen, Y. Yang, chenyang le, X. yun Li, H. Zhang, Y. Wen, W. Zhang, J. Wang, and B. XU · 2022
Later among the works it cites.
Density estimation for conservative q-learning, 2022
P. Daoudi, M. Barlier, L. D. Santos, and A. Virmaux · 2022
Later among the works it cites.
Robust imitation learning from corrupted demonstrations, 2022
L. Liu, Z. Tang, L. Li, and D. Luo · 2022
Later among the works it cites.
When should we prefer offline reinforcement learning over behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Later among the works it cites.
Offline rl with realistic datasets: Heteroskedasticity and support constraints
A. Singh, A. Kumar, Q. Vuong, Y. Chebotar, and S. Levine · 2022
Later among the works it cites.
DASCO: Dual-generator adversarial support constrained offline reinforcement learning
Q. Vuong, A. Kumar, S. Levine, and Y. Chebotar · 2022
Later among the works it cites.
Distance-sensitive offline reinforcement learning
J. Li, X. Zhan, H. Xu, X. Zhu, J. Liu, and Y.-Q. Zhang · 2022
Later among the works it cites.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al · 2022
Later among the works it cites.
When does return-conditioned supervised learning work for offline reinforcement learning?
D. Brandfonbrener, A. Bietti, J. Buckman, R. Laroche, and J. Bruna · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl, 2023
T. Yamagata, A. Khalil, and R. Santos-Rodriguez · 2023
Closest in time.