Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), it is easier to solve a task if given a good representation.
Self-supervised learning of image embedding for continuous control
Florensa, C., Degrave, J., Heess, N., Springenberg, J. T., and Riedmiller, M. (2019) · 1901
Earlier work this paper cites.
Towards characterizing divergence in deep Q-learning
Achiam, J., Knight, E., and Abbeel, P. (2019) · 1903
Earlier work this paper cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S. (2019) · 1903
Earlier work this paper cites.
Reinforcement learning without ground-truth state
Lin, X., Baweja, H. S., and Held, D. (2019) · 1905
Earlier work this paper cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V. (2019) · 1906
Earlier work this paper cites.
On mutual information maximization for representation learning
Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S., and Lucic, M. (2019) · 1907
Earlier work this paper cites.
Self-supervised learning of distance functions for goal-conditioned reinforcement learning
Venkattaramanujam, S., Crawford, E., Doan, T. V., and Precup, D. (2019) · 1907
Earlier work this paper cites.
Task-relevant adversarial imitation learning
Zolna, K., Reed, S., Novikov, A., Colmenarejo, S. G., Budden, D., Cabi, S., Denil, M., de Freitas, N., and Wang, Z. (2019) · 1910
Earlier work this paper cites.
Positive-unlabeled reward learning
Xu, D. and Denil, M. (2019) · 1911
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2019a) · 1912
Earlier work this paper cites.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J. (2019) · 1912
Earlier work this paper cites.
Variational empowerment as representation learning for goal-conditioned reinforcement learning
Choi, J., Sharma, A., Lee, H., Levine, S., and Gu, S. S. (2021) · 1963
Earlier work this paper cites.
Survey sampling
Kish, L. (1965) · 1965
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. (1993) · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P. (1993) · 1993
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J. (1999) · 1999
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E. (2020) · 2002
Earlier work this paper cites.
Rewriting history with inverse RL: Hindsight inference for policy improvement
Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. (2020) · 2002
Earlier work this paper cites.
D4RL: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Weinberger, K. Q. and Saul, L. K. (2005) · 2005
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altch’e, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. Á., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. (2020) · 2006
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
Hoffman, M., Shahriari, B., Aslanides, J., Barth-Maron, G., Behbahani, F., Norman, T., Abdolmaleki, A., Cassirer, A., Yang, F., Baumli, K., Henderson, S., Novikov, A., Colmenarejo, S. G., Cabi, S., Gulcehre, C., Paine, T. L., Cowie, A., Wang, Z., Piot, B., and de Freitas, N. (2020) · 2006
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Wilson, A., Fern, A., Ray, S., and Tadepalli, P. (2007) · 2007
Earlier work this paper cites.
Broadly-exploring, local-policy trees for long-horizon task planning
Ichter, B., Sermanet, P., and Lynch, C. (2020) · 2010
Earlier work this paper cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S. (2020) · 2010
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Lange, S. and Riedmiller, M. (2010) · 2010
Earlier work this paper cites.
Specializations of the master problem
Langford, J. (2010) · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D. (2010) · 2010
Earlier work this paper cites.
C-learning: Learning to achieve goals via recursive classification
Eysenbach, B., Salakhutdinov, R., and Levine, S. (2021b) · 2011
Earlier work this paper cites.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Gutmann, M. U. and Hyvärinen, A. (2012) · 2012
Earlier work this paper cites.
Semi-supervised reward learning for offline reinforcement learning
Konyushkova, K., Zolna, K., Aytar, Y., Novikov, A., Reed, S., Cabi, S., and de Freitas, N. (2020) · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, A. and Teh, Y. W. (2012) · 2012
Earlier work this paper cites.
Planning from pixels using inverse dynamics models
Paster, K., McIlraith, S. A., and Ba, J. (2020) · 2012
Earlier work this paper cites.
Learning grasps for unknown objects in cluttered scenes
Fischinger, D., Vincze, M., and Jiang, Y. (2013) · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Levy, O. and Goldberg, Y. (2014) · 2014
Earlier work this paper cites.
Deep metric learning using triplet network
Hoffer, E. and Ailon, N. (2015) · 2015
Earlier work this paper cites.
State of the art control of Atari games using shallow reinforcement learning
Liang, Y., Machado, M. C., Talvitie, E., and Bowling, M. (2015) · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J. (2015) · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. (2015) · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Learning to act by predicting the future
Dosovitskiy, A. and Koltun, V. (2016) · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S. (2016) · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y. (2016) · 2016
Cited alongside, same era.
Raster fairy
Klingemann, M. (2016) · 2016
Cited alongside, same era.
Predictive learning
LeCun, Y. (2016) · 2016
Cited alongside, same era.
On variational bounds of mutual information
Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G. (2019) · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K. (2019) · 2019
Later among the works it cites.
Policy continuation with hindsight inverse dynamics
Sun, H., Li, Z., Liu, X., Zhou, B., and Lin, D. (2019) · 2019
Later among the works it cites.
Solar: Deep structured representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S. (2019) · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
Zhao, R., Sun, X., and Tresp, V. (2019) · 2019
Later among the works it cites.
Learning to reach goals via iterated supervised learning
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C. M., Eysenbach, B., and Levine, S. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
f-GAN: Training generative neural samplers using variational divergence minimization
Nowozin, S., Cseke, B., and Tomioka, R. (2016) · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K. (2016) · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S. (2017) · 2017
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. (2017) · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Guo, Z. D., Pires, B. A., Piot, B., Grill, J.-B., Altché, F., Munos, R., and Azar, M. G. (2020) · 2020
Later among the works it cites.
Self-supervised co-training for video representation learning
Han, T., Xie, W., and Zisserman, A. (2020) · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Later among the works it cites.
gamma-models: Generative temporal difference learning for infinite-horizon prediction
Janner, M., Mordatch, I., and Levine, S. (2020) · 2020
Later among the works it cites.
Generalized hindsight for reinforcement learning
Li, A., Pinto, L., and Abbeel, P. (2020) · 2020
Later among the works it cites.
Hallucinative topological memory for zero-shot visual planning
Liu, K., Kurutach, T., Tung, C., Abbeel, P., and Tamar, A. (2020) · 2020
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P. (2020) · 2020
Later among the works it cites.
Learning predictive models from observation and interaction
Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C. (2020) · 2020
Later among the works it cites.
Predictive coding for locally-linear control
Shu, R., Nguyen, T., Chow, Y., Pham, T., Than, K., Ghavamzadeh, M., Ermon, S., and Bui, H. (2020) · 2020
Later among the works it cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P. (2020) · 2020
Later among the works it cites.
Neural methods for point-wise dependency estimation
Tsai, Y.-H., Zhao, H., Yamada, M., Morency, L.-P., and Salakhutdinov, R. (2020) · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R. (2020) · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
Blier, L., Tallec, C., and Ollivier, Y. (2021) · 2021
Later among the works it cites.
Goal-conditioned reinforcement learning with imagined subgoals
Chane-Sane, E., Schmid, C., and Laptev, I. (2021) · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Chen, X. and He, K. (2021) · 2021
Later among the works it cites.
Curious representation learning for embodied intelligence
Du, Y., Gan, C., and Isola, P. (2021) · 2021
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?
Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. (2021) · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S. (2021) · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S. (2021) · 2021
Later among the works it cites.
CIC: Contrastive intrinsic control for unsupervised skill discovery
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P. (2021) · 2021
Later among the works it cites.
Aps: Active pretraining with successor features
Liu, H. and Abbeel, P. (2021) · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D. (2021) · 2021
Later among the works it cites.
Which mutual-information representation learning objectives are sufficient for control?
Rakelly, K., Gupta, A., Florensa, C., and Levine, S. (2021) · 2021
Later among the works it cites.
Outcome-driven reinforcement learning via variational inference
Rudner, T. G., Pong, V., McAllister, R., Gal, Y., and Levine, S. (2021) · 2021
Later among the works it cites.
Model-based reinforcement learning via latent-space collocation
Rybkin, O., Zhu, C., Nagabandi, A., Daniilidis, K., Mordatch, I., and Levine, S. (2021) · 2021
Later among the works it cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021) · 2021
Later among the works it cites.
Unsupervised learning for reinforcement learning
Srinivas, A. and Abbeel, P. (2021) · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M. (2021) · 2021
Later among the works it cites.
Overcoming the spectral bias of neural value approximation
Yang, G., Ajay, A., and Agrawal, P. (2021) · 2021
Later among the works it cites.
Imitating past successes can be very suboptimal
Eysenbach, B., Udatha, S., Levine, S., and Salakhutdinov, R. (2022) · 2022
Closest in time.
Hong, Z.-W., Yang, G., and Agrawal, P. (2022) · 2022
Closest in time.
Private Communication
IQL · 2022
Closest in time.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
Nair, S., Mitchell, E., Chen, K., Savarese, S., Finn, C., et al. (2022) · 2022
Closest in time.
Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning
Qiu, S., Wang, L., Bai, C., Yang, Z., and Wang, Z. (2022) · 2022
Closest in time.
Investigating the properties of neural network representations in reinforcement learning
Wang, H., Miahi, E., White, M., Machado, M. C., Abbas, Z., Kumaraswamy, R., Liu, V., and White, A. (2022) · 2022
Closest in time.
Making linear mdps practical via contrastive representation learning
Zhang, T., Ren, T., Yang, M., Gonzalez, J., Schuurmans, D., and Dai, B. (2022) · 2022
Closest in time.