Fetching the paper…
Reading the bibliography…
We are interested in learning scalable agents for reinforcement learning that can learn from large-scale, diverse sequential data similar to current large vision and language models.
An intrinsic reward mechanism for efficient exploration
Ö. Şimşek and A. G. Barto · 2006
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
P.-Y. Oudeyer and F. Kaplan · 2009
Earlier work this paper cites.
Intrinsically motivated learning in natural and artificial systems
G. Baldassarre and M. Mirolli · 2013
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
P. Christiano, Z. Shah, I. Mordatch, J. Schneider, T. Blackwell, J. Tobin, P. Abbeel, and W. Zaremba · 2016
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Earlier work this paper cites.
Deepmind control suite
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
Earlier work this paper cites.
Behavioral cloning from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Earlier work this paper cites.
Scene memory transformer for embodied agents in long-horizon tasks
K. Fang, A. Toshev, L. Fei-Fei, and S. Savarese · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Cited alongside, same era.
Vd-bert: A unified vision and dialog transformer with bert
Y. Wang, S. Joty, M. R. Lyu, I. King, C. Xiong, and S. C. Hoi · 2020
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, R. D. Hjelm, P. Bachman, and A. C. Courville · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2021
Later among the works it cites.
Training data-efficient image transformers and distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A framework for efficient robotic manipulation
A. Zhan, P. Zhao, L. Pinto, P. Abbeel, and M. Laskin · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Cited alongside, same era.
Urlb: Unsupervised reinforcement learning benchmark
M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Cited alongside, same era.
Uni[mask]: Unified inference in sequential decision problems
M. Carroll, J. Lin, O. Paradise, R. Georgescu, M. Sun, D. Bignell, S. Milani, K. Hofmann, M. Hausknecht, A. Dragan, and S. Devlin · 2022
Closest in time.
Generative pretraining for black-box optimization
S. Krishnamoorthy, S. M. Mashkaria, and A. Grover · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
M. Laskin, H. Liu, X. B. Peng, D. Yarats, A. Rajeswaran, and P. Abbeel · 2022
Closest in time.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling
T. Nguyen and A. Grover · 2022
Closest in time.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Closest in time.
Q. Zheng, A. Zhang, and A. Grover · 2022
Closest in time.