Fetching the paper…
Reading the bibliography…
Passive observational data, such as human videos, is abundant and rich in information, yet remains largely untapped by current RL methods.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M. Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J. Hunt, Tom Schaul, David Silver, and H. V. Hasselt · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
Tejas D. Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J. Gershman · 2016
Earlier work this paper cites.
Eigenoption discovery through the deep successor representation
Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell · 2017
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, and Sergey Levine · 2017
Earlier work this paper cites.
Third-person imitation learning
Bradly C. Stadie, P. Abbeel, and Ilya Sutskever · 2017
Earlier work this paper cites.
Universal successor features approximators
Diana Borsa, André Barreto, John Quan, Daniel Jaymin Mankowitz, Rémi Munos, H. V. Hasselt, David Silver, and Tom Schaul · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
Ashvin Nair, Vitchyr H. Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Perceptual values from observation
Ashley D. Edwards and Charles Lee Isbell · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, D. Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, G. Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Learning one representation to optimize all rewards
Ahmed Touati and Yann Ollivier · 2021
Later among the works it cites.
Xirl: Cross-embodiment inverse reinforcement learning
Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Later among the works it cites.
Learning value functions from undirected state-only experience
Matthew Chang, Arjun Gupta, and Saurabh Gupta · 2022
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Reinforcement learning with videos: Combining offline observations with interaction
Karl Schmeckpeper, Oleh Rybkin, Kostas Daniilidis, Sergey Levine, and Chelsea Finn · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, Michael Laskin, and P. Abbeel · 2020
Cited alongside, same era.
The MAGICAL benchmark for robust imitation
Sam Toyer, Rohin Shah, Andrew Critch, and Stuart Russell · 2020
Cited alongside, same era.
Ego4d: Around the world in 3, 000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Christian Fuegen, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Cited alongside, same era.
Masked world models for visual control
Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and P. Abbeel
Cited in the paper.
Benjamin Eysenbach, Tianjun Zhang, Ruslan Salakhutdinov, and Sergey Levine · 2022
Later among the works it cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhi Gupta · 2022
Later among the works it cites.
You can’t count on luck: Why decision transformers fail in stochastic environments
Keiran Paster, Sheila A. McIlraith, and Jimmy Ba · 2022
Later among the works it cites.
Addressing optimism bias in sequence modeling for reinforcement learning
Adam R Villaflor, Zhe Huang, Swapnil Pande, John M Dolan, and Jeff Schneider · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik · 2022
Later among the works it cites.
Dichotomy of control: Separating what you can control from what you cannot
Mengjiao Yang, Dale Schuurmans, P. Abbeel, and Ofir Nachum · 2022
Later among the works it cites.