Fetching the paper…
Reading the bibliography…
Skills or low-level policies in reinforcement learning are temporally extended actions that can speed up learning and enable complex behaviours.
A bayesian analysis of some nonparametric problems
Ferguson, T. S · 1973
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B · 1977
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Glynn, P. W · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A constructive definition of dirichlet priors
Sethuraman, J · 1994
Earlier work this paper cites.
Behavioural cloning in control of a dynamic system
Esmaili, N., Sammut, C., and Shirazi, G · 1995
Earlier work this paper cites.
Python reference manual
vanRossum, G · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning from demonstration
Schaal, S. et al · 1997
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
Dietterich, T. G. et al · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Markov chain sampling methods for dirichlet process mixture models
Neal, R. M · 2000
Earlier work this paper cites.
Gibbs sampling methods for stick-breaking priors
Ishwaran, H. and James, L. F · 2001
Earlier work this paper cites.
Hierarchical dirichlet processes
Teh, Y. W., Jordan, M. I., Beal, M. J., and Blei, D. M · 2006
Earlier work this paper cites.
Generalized polya urn for time-varying dirichlet process mixtures
Caron, F., Davy, M., and Doucet, A · 2007
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E · 2007
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Nonparametric bayes modeling of multivariate categorical data
Dunson, D. B. and Xing, C · 2009
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Koller, D. and Friedman, N · 2009
Earlier work this paper cites.
Revisiting k-means: New algorithms via bayesian nonparametrics
Kulis, B. and Jordan, M. I · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Small-variance asymptotics for exponential family dirichlet process mixture models
Jiang, K., Kulis, B., and Jordan, M · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Mad-bayes: Map-based asymptotic derivations from bayes
Broderick, T., Kulis, B., and Jordan, M · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
Incremental semantically grounded learning from demonstration
Niekum, S., Chitta, S., Barto, A. G., Marthi, B., and Osentoski, S · 2013
Cited alongside, same era.
Small-variance asymptotics for hidden markov models
Roychowdhury, A., Jiang, K., and Kulis, B · 2013
Cited alongside, same era.
Amortized inference in probabilistic reasoning
Gershman, S. and Goodman, N · 2014
Sticking the landing: Simple, lower-variance gradient estimators for variational inference
Roeder, G., Wu, Y., and Duvenaud, D. K · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Importance weighted autoencoders
Burda, Y., Grosse, R., and Salakhutdinov, R · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Later among the works it cites.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning
Krishnan, S., Garg, A., Patil, S., Lea, C., Hager, G., Abbeel, P., and Goldberg, K · 2018
Later among the works it cites.
Sharma, A., Sharma, M., Rhinehart, N., and Kitani, K. M · 2018
Later among the works it cites.
TACO: Learning task decomposition via temporal alignment for control
Shiarlis, K., Wulfmeier, M., Salter, S., Whiteson, S., and Posner, I · 2018
Later among the works it cites.
Compile: Compositional imitation learning and execution
Kipf, T., Li, Y., Dai, H., Zambaldi, V., Sanchez-Gonzalez, A., Grefenstette, E., Kohli, P., and Battaglia, P · 2019
Later among the works it cites.
Swirl: A sequential windowed inverse reinforcement learning algorithm for robot tasks with delayed rewards
Krishnan, S., Garg, A., Liaw, R., Thananjeyan, B., Miller, L., Pokorny, F. T., and Goldberg, K · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Later among the works it cites.
A new distribution on the simplex with auto-encoding applications
Stirn, A., Jebara, T., and Knowles, D · 2019
Later among the works it cites.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Such, F. P., Madhavan, V., Liu, R., Wang, R., Castro, P. S., Li, Y., Zhi, J., Schubert, L., Bellemare, M. G., Clune, J., et al · 2019
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
Options of interest: Temporal abstraction with interest functions
Khetarpal, K., Klissarov, M., Chevalier-Boisvert, M., Bacon, P.-L., and Precup, D · 2020
Later among the works it cites.
Estimating gradients for discrete random variables by sampling without replacement
Kool, W., van Hoof, H., and Welling, M · 2020
Later among the works it cites.
Rao-blackwellizing the straight-through gumbel-softmax gradient estimator
Paulus, M. B., Maddison, C. J., and Krause, A · 2020
Later among the works it cites.
Invertible gaussian reparameterization: Revisiting the gumbel-softmax
Potapczynski, A., Loaiza-Ganem, G., and Cunningham, J. P · 2020
Later among the works it cites.
Learning robot skills with temporal variational inference
Shankar, T. and Gupta, A · 2020
Later among the works it cites.
{OPAL}: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Later among the works it cites.