Fetching the paper…
Reading the bibliography…
Identifying statistical regularities in solutions to some tasks in multi-task reinforcement learning can accelerate the learning of new tasks.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Univresal sequential coding of single messages
Yuri M Shtarkov · 1987
Earlier work this paper cites.
Autoencoders, minimum description length, and helmholtz free energy
Geoffrey E Hinton and Richard S Zemel · 1994
Earlier work this paper cites.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1994
Earlier work this paper cites.
Macro-actions in reinforcement learning: An empirical analysis
Amy McGovern and Richard S Sutton · 1998
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Variational learning for switching state-space models
Zoubin Ghahramani and Geoffrey E Hinton · 2000
Earlier work this paper cites.
Topic segmentation with an aspect hidden markov model
David M Blei and Pedro J Moreno · 2001
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2003
Earlier work this paper cites.
A tutorial introduction to the minimum description length principle
Peter Grunwald · 2004
Earlier work this paper cites.
Variational learning and bits-back coding: an information-theoretic view to bayesian learning
Antti Honkela and Harri Valpola · 2004
Earlier work this paper cites.
Self-paced learning for latent variable models
M Pawan Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
A sticky hdp-hmm with application to speaker diarization
Emily B Fox, Erik B Sudderth, Michael I Jordan, and Alan S Willsky · 2011
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Incremental semantically grounded learning from demonstration
Scott Niekum, Sachin Chitta, Andrew G Barto, Bhaskara Marthi, and Sarah Osentoski · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Earlier work this paper cites.
Recurrent hidden semi-markov model
Hanjun Dai, Bo Dai, Yan-Ming Zhang, Shuang Li, and Le Song · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Composing graphical models with neural networks for structured representations and fast inference
Matthew J Johnson, David K Duvenaud, Alex Wiltschko, Ryan P Adams, and Sandeep R Datta · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
Recurrent switching linear dynamical systems
Scott W Linderman, Andrew C Miller, Ryan P Adams, David M Blei, Liam Paninski, and Matthew J Johnson · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Continuous meta-learning without tasks
James Harrison, Apoorva Sharma, Chelsea Finn, and Marco Pavone · 2019
Later among the works it cites.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Gu, Kevin Murphy, and Chelsea Finn · 2019
Later among the works it cites.
Variational temporal abstraction
Taesup Kim, Sungjin Ahn, and Yoshua Bengio · 2019
Later among the works it cites.
Compile: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Multi-level discovery of deep options
Roy Fox, Sanjay Krishnan, Ion Stoica, and Ken Goldberg · 2017
Cited alongside, same era.
A compression-inspired framework for macro discovery
Francisco M Garcia, Bruno C da Silva, and Philip S Thomas · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Dhruva TB, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, SM Eslami, et al · 2017
Cited alongside, same era.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Sanjay Krishnan, Roy Fox, Ion Stoica, and Ken Goldberg · 2017
Cited alongside, same era.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2019
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2020
Later among the works it cites.
Collapsed amortized variational inference for switching nonlinear dynamical systems
Zhe Dong, Bryan Seybold, Kevin Murphy, and Hung Bui · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang*, Behnam Neyshabur*, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Learning compound tasks without task-specific knowledge via imitation and self-supervised learning
Sang-Hyun Lee and Seung-Woo Seo · 2020
Later among the works it cites.
Learning abstract models for strategic exploration and fast reward transfer
Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri, Kelvin Guu, Panupong Pasupat, Emma Brunskill, and Percy Liang · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
Karl Pertsch, Youngwoon Lee, and Joseph J Lim · 2020
Later among the works it cites.
Learning robot skills with temporal variational inference
Tanmay Shankar and Abhinav Gupta · 2020
Later among the works it cites.
Discovering motor programs by recomposing demonstrations
Tanmay Shankar, Shubham Tulsiani, Lerrel Pinto, and Abhinav Gupta · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Augmenting policy learning with routines discovered from a demonstration
Zelin Zhao, Chuang Gan, Jiajun Wu, Xiaoxiao Guo, and Joshua B Tenenbaum · 2020
Later among the works it cites.
Plas: Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2020
Later among the works it cites.
Skill discovery for exploration and planning using deep skill graphs
Akhil Bagaria, Jason K Senthil, and George Konidaris · 2021
Later among the works it cites.
Learning task decomposition with ordered memory policy network
Yuchen Lu, Yikang Shen, Siyuan Zhou, Aaron Courville, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Skill-based meta-reinforcement learning
Taewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang, and Joseph J Lim · 2021
Later among the works it cites.
Learning transferable motor skills with hierarchical latent mixture policies
Dushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier, Martina Zambelli, Giulia Vezzani, Dhruva Tirumala, Yusuf Aytar, Josh Merel, Nicolas Heess, et al · 2021
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2021
Later among the works it cites.
Sumedh A Sontakke, Sumegh Roychowdhury, Mausoom Sarkar, Nikaash Puri, Balaji Krishnamurthy, and Laurent Itti · 2021
Later among the works it cites.
Skid raw: Skill discovery from raw trajectories
Daniel Tanneberg, Kai Ploeger, Elmar Rueckert, and Jan Peters · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Mengjiao Yang and Ofir Nachum · 2021
Later among the works it cites.
Trail: Near-optimal imitation learning with suboptimal data
Mengjiao Yang, Sergey Levine, and Ofir Nachum · 2021
Later among the works it cites.
Minimum description length skills for accelerated reinforcement learning
Jesse Zhang, Karl Pertsch, Jiefan Yang, and Joseph J Lim · 2021
Later among the works it cites.
Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation
Yifeng Zhu, Peter Stone, and Yuke Zhu · 2021
Later among the works it cites.