Fetching the paper…
Reading the bibliography…
Unsupervised pre-training has recently become the bedrock for computer vision and natural language processing.
Learning to generate sub-goals for action sequences
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1992
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Fitted q-iteration by advantage weighted regression
Gerhard Neumann and Jan Peters · 2008
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Muelling, and Yasemin Altun · 2010
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
G. Brockman, Vicki Cheung, Ludwig Pettersson, J. Schneider, John Schulman, Jie Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Joshua B. Tenenbaum · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Sanjay Krishnan, Roy Fox, Ion Stoica, and Ken Goldberg · 2017
Earlier work this paper cites.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Manfred Otto Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Earlier work this paper cites.
Temporal difference models: Model-free deep rl for model-based control
Vitchyr H. Pong, Shixiang Shane Gu, Murtaza Dalal, and Sergey Levine · 2018
Earlier work this paper cites.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Goal-conditioned imitation learning
Yiming Ding, Carlos Florensa, Mariano Phielipp, and P. Abbeel · 2019
Earlier work this paper cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2019
Earlier work this paper cites.
Dher: Hindsight experience replay for dynamic goals
Meng Fang, Cheng Zhou, Bei Shi, Boqing Gong, Jia Xu, and Tong Zhang · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Cited alongside, same era.
Mapping state space using landmarks for universal goal reaching
Zhiao Huang, Fangchen Liu, and Hao Su · 2019
Cited alongside, same era.
Learning multi-level hierarchies with hindsight
Andrew Levy, George Dimitri Konidaris, Robert W. Platt, and Kate Saenko · 2019
Cited alongside, same era.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2019
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2019
Cited alongside, same era.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr H. Pong, Steven Lin, and Sergey Levine · 2019
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2021
Later among the works it cites.
Recon: Rapid exploration for open-world navigation with latent goal models
Dhruv Shah, Benjamin Eysenbach, Gregory Kahn, Nicholas Rhinehart, and Sergey Levine · 2021
Later among the works it cites.
Data-efficient hindsight off-policy option learning
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Siegel, Nicolas Manfred Otto Heess, and Martin A. Riedmiller · 2021
Later among the works it cites.
World model as a graph: Learning latent landmarks for planning
Lunjun Zhang, Ge Yang, and Bradly C. Stadie · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, G. Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, G. Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Learning value functions from undirected state-only experience
Matthew Chang, Arjun Gupta, and Saurabh Gupta · 2022
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
Benjamin Eysenbach, Tianjun Zhang, Ruslan Salakhutdinov, and Sergey Levine · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll’ar, and Ross B. Girshick · 2022
Later among the works it cites.
Bilinear value networks
Zhang-Wei Hong, Ge Yang, and Pulkit Agrawal · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
Dr3: Value-based deep reinforcement learning requires explicit regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville, G. Tucker, and Sergey Levine · 2022
Later among the works it cites.
Hierarchical planning through goal-conditioned offline reinforcement learning
Jinning Li, Chen Tang, Masayoshi Tomizuka, and Wei Zhan · 2022
Later among the works it cites.
How far i’ll go: Offline goal-conditioned reinforcement learning via f-advantage regression
Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhi Gupta · 2022
Later among the works it cites.
Latent plans for task-agnostic offline reinforcement learning
Erick Rosete-Beas, Oier Mees, Gabriel Kalweit, Joschka Boedecker, and Wolfram Burgard · 2022
Later among the works it cites.
Mo2: Model-based offline options
Sasha Salter, Markus Wulfmeier, Dhruva Tirumala, Nicolas Manfred Otto Heess, Martin A. Riedmiller, Raia Hadsell, and Dushyant Rao · 2022
Later among the works it cites.
Reinforcement learning with action-free pre-training from videos
Younggyo Seo, Kimin Lee, Stephen James, and P. Abbeel · 2022
Later among the works it cites.
Skill-based model-based reinforcement learning
Lu Shi, Joseph J. Lim, and Youngwoon Lee · 2022
Later among the works it cites.
Addressing optimism bias in sequence modeling for reinforcement learning
Adam R. Villaflor, Zheng Huang, Swapnil Pande, John M. Dolan, and Jeff G. Schneider · 2022
Later among the works it cites.
A policy-guided imitation approach for offline reinforcement learning
Haoran Xu, Li Jiang, Jianxiong Li, and Xianyuan Zhan · 2022
Later among the works it cites.
Rethinking goal-conditioned supervised learning and its connection to offline rl
Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, and Chongjie Zhang · 2022
Later among the works it cites.
C-planning: An automatic curriculum for learning goal-reaching tasks
Tianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine, and Joseph Gonzalez · 2022
Later among the works it cites.
Extreme q-learning: Maxent rl without entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon · 2023
Closest in time.
dibyaghosh/jaxrl_m, 2023
Dibya Ghosh · 2023
Closest in time.
Reinforcement learning from passive data via latent intentions
Dibya Ghosh, Chethan Bhateja, and Sergey Levine · 2023
Closest in time.
Efficient planning in a compact latent action space
Zhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li, Tim Rocktaschel, Edward Grefenstette, and Yuandong Tian · 2023
Closest in time.
Imitating graph-based planning with goal-conditioned policies
Junsu Kim, Younggyo Seo, Sungsoo Ahn, Kyunghwan Son, and Jinwoo Shin · 2023
Closest in time.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2023
Closest in time.
Optimal goal-reaching reinforcement learning via quasimetric learning
Tongzhou Wang, Antonio Torralba, Phillip Isola, and Amy Zhang · 2023
Closest in time.
Offline rl with no ood actions: In-sample learning via implicit value regularization
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Chan, and Xianyuan Zhan · 2023
Closest in time.
Dichotomy of control: Separating what you can control from what you cannot
Mengjiao Yang, Dale Schuurmans, P. Abbeel, and Ofir Nachum · 2023
Closest in time.