Fetching the paper…
Reading the bibliography…
Goal-conditioned planning benefits from learned low-dimensional representations of rich observations.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
A note on two problems in connexion with graphs
Dijkstra, E. W · 1959
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Rubinstein, R · 1999
Earlier work this paper cites.
A tutorial on the cross-entropy method
De Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y · 2005
Earlier work this paper cites.
Receding horizon control
Mattingley, J., Wang, Y., and Boyd, S · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos, 2015
Wang, X. and Gupta, A · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Johnson, A. E., Pollard, T. J., Shen, L., Lehman, L.-w. H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Anthony Celi, L., and Mark, R. G · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S · 2017
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning, corr abs/1710.02298
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., and Silver, D · 2017
Earlier work this paper cites.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M · 2017
Earlier work this paper cites.
1 year, 1000 km: The oxford robotcar dataset
Maddern, W., Pascoe, G., Linegar, C., and Newman, P · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Cited alongside, same era.
Dher: Hindsight experience replay for dynamic goals
Fang, M., Zhou, C., Shi, B., Gong, B., Xu, J., and Zhang, T · 2018
Cited alongside, same era.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Cited alongside, same era.
Planning from pixels using inverse dynamics models
Paster, K., McIlraith, S. A., and Ba, J · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Later among the works it cites.
Model-based visual planning with self-supervised functional distances
Tian, S., Nair, S., Ebert, F., Dasari, S., Eysenbach, B., Finn, C., and Levine, S · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K · 2018
Cited alongside, same era.
Bdd100k: A diverse driving video database with scalable annotation tooling
Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V., Darrell, T., et al · 2018
Cited alongside, same era.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M., Dabney, W., Dadashi, R., Ali Taiga, A., Castro, P. S., Le Roux, N., Schuurmans, D., Lattimore, T., and Lyle, C · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J · 2019
Cited alongside, same era.
Probability: theory and examples , volume 49
Durrett, R · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Cited alongside, same era.
Amin, S., Gomrokchi, M., Satija, H., van Hoof, H., and Precup, D · 2021
Later among the works it cites.
Provably filtering exogenous distractors using multistep inverse dynamics
Efroni, Y., Misra, D., Krishnamurthy, A., Agarwal, A., and Langford, J · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W · 2021
Later among the works it cites.
Sample-efficient cross-entropy method for real-time planning
Pinneri, C., Sawant, S., Blaes, S., Achterhold, J., Stueckler, J., Rolinek, M., and Martius, G · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Later among the works it cites.
Provably filtering exogenous distractors using multistep inverse dynamics
Efroni, Y., Misra, D., Krishnamurthy, A., Agarwal, A., and Langford, J · 2022
Later among the works it cites.
Agent-controller representations: Principled offline rl with rich exogenous information
Islam, R., Tomar, M., Lamb, A., Efroni, Y., Zang, H., Didolkar, A., Misra, D., Li, X., van Seijen, H., Combes, R. T. d., et al · 2022
Later among the works it cites.
Guaranteed discovery of controllable latent states with multi-step inverse models
Lamb, A., Islam, R., Efroni, Y., Didolkar, A., Misra, D., Foster, D., Molu, L., Chari, R., Krishnamurthy, A., and Langford, J · 2022
Later among the works it cites.
On the generalization of representations in reinforcement learning
Lan, C. L., Tu, S., Oberman, A., Agarwal, R., and Bellemare, M. G · 2022
Later among the works it cites.
Challenges and opportunities in offline reinforcement learning from visual observations
Lu, C., Ball, P. J., Rudner, T. G. J., Parker-Holder, J., Osborne, M. A., and Teh, Y. W · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation, 2022
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Later among the works it cites.
Bootstrapped representations in reinforcement learning
Lan, C. L., Tu, S., Rowland, M., Harutyunyan, A., Agarwal, R., Bellemare, M. G., and Dabney, W · 2023
Closest in time.
Learning goal-conditioned policies offline with self-supervised reward shaping
Mezghani, L., Sukhbaatar, S., Bojanowski, P., Lazaric, A., and Karteek, A · 2023
Closest in time.
Mhammedi, Z., Foster, D. J., and Rakhlin, A · 2023
Closest in time.
Cr-vae: Contrastive regularization on variational autoencoders for preventing posterior collapse
Rueckert, F. L. et al · 2023
Closest in time.
Optimal goal-reaching reinforcement learning via quasimetric learning
Wang, T., Torralba, A., Isola, P., and Zhang, A · 2023
Closest in time.
Agent-centric state discovery for finite-memory POMDPs
Wu, L., Evans, B., Islam, R., Seraj, R., Efroni, Y., and Lamb, A · 2023
Closest in time.