Fetching the paper…
Reading the bibliography…
Exploration in environments which differ across episodes has received increasing attention in recent years.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, E. L. and Littman, M. L · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y · 2009
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M., Naddaf, Y., Veness, J., and Bowling, M · 2012
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A., Di Castro, D., and Mannor, S · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D., Unterthiner, T., and Hochreiter, S · 2016
Earlier work this paper cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, J. and Sastry, S · 2017
Earlier work this paper cites.
Exploration-exploitation in mdps with options
Fruit, R. and Lazaric, A · 2017
Earlier work this paper cites.
Count-based exploration in feature space for reinforcement learning
Martin, J., Sasikumar, S. N., Everitt, T., and Hutter, M · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Xi Chen, O., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Earlier work this paper cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Generalization and regularization in DQN
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Procedural level generation improves generality of deep reinforcement learning
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D · 2018
Cited alongside, same era.
igibson 1.0: A simulation environment for interactive tasks in large realistic scenes
Shen, B., Xia, F., Li, C., Martín-Martín, R., Fan, L., Wang, G., Pérez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L., et al · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2020
Later among the works it cites.
No-regret exploration in goal-oriented reinforcement learning
Tarbouriech, J., Garcelon, E., Valko, M., Pirotta, M., and Lazaric, A · 2020
Later among the works it cites.
Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A. S., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D · 2020
Later among the works it cites.
Sapien: A simulated part-based interactive environment
Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanton, C. and Clune, J · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Cited alongside, same era.
Explicit explore-exploit algorithms in continuous state spaces
Henaff, M · 2019
Cited alongside, same era.
Obstacle tower: A generalization challenge in vision, control, and planning
Juliani, A., Khalifa, A., Berges, V., Harper, J., Henry, H., Crespi, A., Togelius, J., and Lange, D · 2019
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., and Batra, D · 2019
Cited alongside, same era.
Model-based active exploration
Shyam, P., Jaśkowski, W., and Gomez, F · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Agarwal, A., Henaff, M., Kakade, S., and Sun, W · 2020
Cited alongside, same era.
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
Explore and control with adversarial surprise
Fickinger, A., Jaques, N., Parajuli, S., Chang, M., Rhinehart, N., Berseth, G., Russell, S., and Levine, S · 2021
Later among the works it cites.
Adversarially guided actor-critic
Flet-Berliac, Y., Ferret, J., Pietquin, O., Preux, P., and Geist, M · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Later among the works it cites.
Model-based reinforcement learning with ensembled model-value expansion
Manek, G. and Kolter, J. Z · 2021
Later among the works it cites.
Interesting object, curious agent: Learning task-agnostic exploration
Parisi, S., Dean, V., Pathak, D., and Gupta, A · 2021
Later among the works it cites.
Megaverse: Simulating embodied agents at one million experiences per second
Petrenko, A., Wijmans, E., Shacklett, B., and Koltun, V · 2021
Later among the works it cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J. M., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., Savva, M., Zhao, Y., and Batra, D · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
Don’t do what doesn’t matter: Intrinsic motivation with action usefulness
Seurin, M., Strub, F., Preux, P., and Pietquin, O · 2021
Later among the works it cites.
Habitat 2.0: Training home assistants to rearrange their habitat
Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D., Maksymets, O., Gokaslan, A., Vondrus, V., Dharur, S., Meier, F., Galuba, W., Chang, A., Kira, Z., Koltun, V., Malik, J., Savva, M., and Batra, D · 2021
Later among the works it cites.
Exploration via elliptical episodic bonuses
Henaff, M., Raileanu, R., Jiang, M., and Rocktäschel, T · 2022
Later among the works it cites.
Leco: Learnable episodic count for task-specific intrinsic reward
Jo, D., Kim, S., Nam, D. W., Kwon, T., Rho, S., Kim, J., and Lee, D · 2022
Later among the works it cites.
Exploring through random curiosity with general value functions
Ramesh, A., Kirsch, L., van Steenkiste, S., and Schmidhuber, J · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F · 2022
Later among the works it cites.
Revisiting intrinsic reward for exploration in procedurally generated environments
Wang, K., Zhou, K., Kang, B., Feng, J., and YAN, S · 2023
Closest in time.