Fetching the paper…
Reading the bibliography…
The varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Causation, Prediction, and Search
Spirtes, P., Glymour, C. N., and Scheines, R · 2000
Earlier work this paper cites.
Dynamic bayesian networks: representation, inference and learning
Murphy, K. P · 2002
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al · 2008
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
DirectLiNGAM: A direct method for learning a linear non-Gaussian structural equation model
Shimizu, S., Inazumi, T., Sogawa, Y., Hyvärinen, A., Kawahara, Y., Washio, T., Hoyer, P. O., and Bollen, K · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M., Lang, T., Toussaint, M., and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Reinforcement learning and causal models
Gershman, S. J · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2017
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft q-learning with mutual-information regularization
Grau-Moya, J., Leibfried, F., and Vrancx, P · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
UCB momentum q-learning: Correcting the bias without forgetting
Ménard, P., Domingues, O. D., Shang, X., and Valko, M · 2021
Later among the works it cites.
Hierarchical reinforcement learning: A comprehensive survey
Pateria, S., Subagdja, B., Tan, A.-h., and Quek, C · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Later among the works it cites.
Dagma: Learning dags via m-matrices and a log-determinant acyclicity characterization
Bello, K., Aragam, B., and Ravikumar, P · 2022
Later among the works it cites.
Generalizing goal-conditioned reinforcement learning with variational causal reasoning
Ding, W., Lin, H., Li, B., and Zhao, D · 2022
Later among the works it cites.
Factored adaptation for non-stationary reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., and Zaremba, W · 2018
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Exploration in reinforcement learning with deep covering options
Jinnai, Y., Park, J. W., Machado, M. C., and Konidaris, G · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Cited alongside, same era.
Maximum entropy-regularized multi-goal reinforcement learning
Zhao, R., Sun, X., and Tresp, V · 2019
Cited alongside, same era.
Feng, F., Huang, B., Zhang, K., and Magliacane, S · 2022
Later among the works it cites.
Leveraging approximate symbolic models for reinforcement learning via skill diversity
Guan, L., Sreedharan, S., and Kambhampati, S · 2022
Later among the works it cites.
Exploration in deep reinforcement learning: A survey
Ladosz, P., Weng, L., Kim, M., and Oh, H · 2022
Later among the works it cites.
Human-ai shared control via policy dissection
Li, Q., Peng, Z., Wu, H., Feng, L., and Zhou, B · 2022
Later among the works it cites.
The importance of non-markovianity in maximum state entropy exploration
Mutti, M., Santi, R. D., and Restelli, M · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P., and Courville, A. C · 2022
Later among the works it cites.
Brain-wide mapping reveals that engrams for a single memory are distributed across multiple brain regions
Roy, D. S., Park, Y.-G., Kim, M. E., Zhang, Y., Ogawa, S. K., DiNapoli, N., Gu, X., Cho, J. H., Choi, H., Kamentsky, L., et al · 2022
Later among the works it cites.
Optimistic curiosity exploration and conservative exploitation with linear reward shaping
Sun, H., Han, L., Yang, R., Ma, X., Guo, J., and Zhou, B · 2022
Later among the works it cites.
Causal reinforcement learning: A survey
Deng, Z., Jiang, J., Long, G., and Zhang, C · 2023
Later among the works it cites.
Hierarchical reinforcement learning integrating with human knowledge for practical robot skill learning in complex multi-stage manipulation
Liu, X., Wang, G., Liu, Z., Liu, Y., Liu, Z., and Huang, P · 2023
Later among the works it cites.
Temporal abstraction in reinforcement learning with the successor representation
Machado, M. C., Barreto, A., Precup, D., and Bowling, M · 2023
Later among the works it cites.
Deep reinforcement learning with plasticity injection
Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Drm: Mastering visual reinforcement learning through dormant ratio minimization
Xu, G., Zheng, R., Liang, Y., Wang, X., Yuan, Z., Ji, T., Luo, Y., Liu, X., Yuan, J., Hua, P., et al · 2023
Later among the works it cites.
CEM: constrained entropy maximization for task-agnostic safe exploration
Yang, Q. and Spaan, M. T. J · 2023
Later among the works it cites.
A survey on causal reinforcement learning
Zeng, Y., Cai, R., Sun, F., Huang, L., and Hao, Z · 2023
Later among the works it cites.
Safety-aware causal representation for trustworthy offline reinforcement learning in autonomous driving
Lin, H., Ding, W., Liu, Z., Niu, Y., Zhu, J., Niu, Y., and Zhao, D · 2024
Closest in time.