Fetching the paper…
Reading the bibliography…
Recent exploration methods have proven to be a recipe for improving sample-efficiency in deep reinforcement learning (RL).
Further experiments with papa
Gamba, A., Gamberini, L., Palmieri, G., and Sanna, R · 1961
Earlier work this paper cites.
Perceptrons: An introduction to computational geometry
Minsky, M. and Papert, S. A · 1969
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Beirlant, J., Dudewicz, E. J., Györfi, L., and Van der Meulen, E. C · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H., Misra, N., Hnizdo, V., Fedorowicz, A., and Demchuk, E · 2003
Earlier work this paper cites.
The random projection method , volume 65
Vempala, S. S · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
On random weights and unsupervised feature learning
Saxe, A. M., Koh, P. W., Chen, Z., Bhand, M., Suresh, B., and Ng, A. Y · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Later among the works it cites.
No training required: Exploring random encoders for sentence classification
Wieting, J. and Kiela, D · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., et al · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Later among the works it cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Later among the works it cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Lee, K., Lee, K., Shin, J., and Lee, H · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Condensenet: An efficient densenet using learned group convolutions
Huang, G., Liu, S., Van der Maaten, L., and Weinberger, K. Q · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Cited alongside, same era.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Cited alongside, same era.
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, A., Laskin, M., and Abbeel, P · 2020
Later among the works it cites.
Benchmarking bonus-based exploration methods on the arcade learning environment
Taïga, A. A., Fedus, W., Machado, M. C., Courville, A., and Bellemare, M. G · 2020
Later among the works it cites.
Novelty search in representational space for sample efficient exploration
Tao, R. Y., François-Lavet, V., and Pineau, J · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N · 2020
Later among the works it cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy, 2010
Ziebart, B. D · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Closest in time.
Behavior from the void: Unsupervised active pre-training
Liu, H. and Abbeel, P · 2021
Closest in time.
A policy gradient method for task-agnostic exploration
Mutti, M., Pratissoli, L., and Restelli, M · 2021
Closest in time.
Learning to plan optimistically: Uncertainty-guided deep exploration via latent model ensembles, 2021
Seyde, T., Schwarting, W., Karaman, S., and Rus, D · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2021
Closest in time.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2021
Closest in time.