Fetching the paper…
Reading the bibliography…
Using massive datasets to train large-scale models has emerged as a dominant approach for broad generalization in natural language and vision applications.
Evolutionary principles in self-referential learning. (on learning how to learn: The meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1991
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Partially labeled classification with markov random walks
Martin Szummer and Tommi Jaakkola · 2001
Earlier work this paper cites.
Algorithm selection via meta-learning
Alexandros Kalousis · 2002
Earlier work this paper cites.
Learning from labeled and unlabeled data with label propagation
Xiaojin Zhu and Zoubin Ghahramani · 2002
Earlier work this paper cites.
Learning inverse dynamics: A comparison
D. Nguyen-Tuong, J. Peters, M. Seeger, and B. Schölkopf · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E. Taylor and Peter Stone · 2009
Earlier work this paper cites.
Introduction to Semi-Supervised Learning
Xiaojin Zhu, Andrew B. Goldberg, Ronald Brachman, and Thomas Dietterich · 2009
Earlier work this paper cites.
Sparse multi-task reinforcement learning
Daniele Calandriello, Alessandro Lazaric, and Marcello Restelli · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Diederik P. Kingma, Danilo J. Rezende, Shakir Mohamed, and Max Welling · 2014
Earlier work this paper cites.
Facial landmark detection by deep multi-task learning
Zhanpeng Zhang, Ping Luo, Chen Change Loy, and Xiaoou Tang · 2014
Earlier work this paper cites.
Semi-supervised learning with ladder networks
Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, and Tapani Raiko · 2015
Earlier work this paper cites.
Learning shared representations in multi-task reinforcement learning, 2016
Diana Borsa, Thore Graepel, and John Shawe-Taylor · 2016
Earlier work this paper cites.
Combining model-based policy search with online model learning for control of physical humanoids
Igor Mordatch, Nikhil Mishra, Clemens Eppner, and Pieter Abbeel · 2016
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh, and Honglak Lee · 2016
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Lei Jimmy Ba, and Ruslan Salakhutdinov · 2016
Cited alongside, same era.
Policy distillation
Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Cited alongside, same era.
RL^2: Fast reinforcement learning via slow reinforcement learning, 2017
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
An overview of multi-task learning in deep neural networks, 2017
Sebastian Ruder · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Sharing knowledge in multi-task deep reinforcement learning
Carlo D’Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli, and Jan Peters · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Adapting Deep Visuomotor Representations with Weak Pairwise Constraints
Eric Tzeng, Coline Devin, Judy Hoffman, Chelsea Finn, Pieter Abbeel, Sergey Levine, Kate Saenko, Trevor Darrell, Pieter Abbeel, Kostas Bekris, and Lauren Miller · 2020
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2020
Yifan Wu, George Tucker, and Ofir Nachum · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio · 2018
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Ignasi Clavera, Anusha Nagabandi, Simin Liu, Ronald S. Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Konrad Zolna, Alexander Novikov, Ksenia Konyushkova, Caglar Gulcehre, Ziyu Wang, Yusuf Aytar, Misha Denil, Nando de Freitas, and Scott Reed · 2020
Later among the works it cites.
A meta-learning approach for automated hyperparameter tuning in evolving data streams
Thomas Lacombe, Yun Sing Koh, Gillian Dobbie, and Ocean Wu · 2021
Later among the works it cites.
Multi-task reinforcement learning with context-based representations
Shagun Sodhani, Amy Zhang, and Joelle Pineau · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Scaling vision transformers, 2021
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos, 2022
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Closest in time.
Multi-game decision transformers, 2022
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, and Igor Mordatch · 2022
Closest in time.
Offline meta-reinforcement learning with online self-supervision
Vitchyr H Pong, Ashvin V Nair, Laura M Smith, Catherine Huang, and Sergey Levine · 2022
Closest in time.
A generalist agent, 2022
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas · 2022
Closest in time.
Reinforcement learning with action-free pre-training from videos
Younggyo Seo, Kimin Lee, Stephen L James, and Pieter Abbeel · 2022
Closest in time.
Chain of thought imitation with procedure cloning
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum · 2022
Closest in time.