Fetching the paper…
Reading the bibliography…
In goal-reaching reinforcement learning (RL), the optimal value function has a particular geometry, called quasimetric structure.
Linear programming and sequential decisions
Manne, A. S · 1960
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B · 1968
Earlier work this paper cites.
On linear programming in a markov decision problem
Denardo, E. V · 1970
Earlier work this paper cites.
Heuristics: intelligent search strategies for computer problem solving
Pearl, J · 1984
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
An abstract approach to dissipation
Sontag, E · 1995
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
Roweis, S. T. and Saul, L. K · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Tenenbaum, J. B., De Silva, V., and Langford, J. C · 2000
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
Proto-value functions: A Laplacian framework for learning representation and control in markov decision processes
Mahadevan, S. and Maggioni, M · 2007
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Bisimulation metrics are optimal value functions
Ferns, N. and Precup, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E · 2015
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
Cited alongside, same era.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M., Dabney, W., Dadashi, R., Ali Taiga, A., Castro, P. S., Le Roux, N., Schuurmans, D., Lattimore, T., and Lyle, C · 2019
Cited alongside, same era.
Learning to reach goals via iterated supervised learning
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C., Eysenbach, B., and Levine, S · 2019
Cited alongside, same era.
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2021
Later among the works it cites.
Learning task informed abstractions
Fu, X., Yang, G., Agrawal, P., and Jaakkola, T · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Instabilities of offline rl with pre-trained neural representation
Wang, R., Wu, Y., Salakhutdinov, R., and Kakade, S · 2021
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Cited alongside, same era.
Kumar, A., Peng, X. B., and Levine, S · 2019
Cited alongside, same era.
Scalable methods for computing state similarity in deterministic markov decision processes
Castro, P. S · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Predictive information accelerates learning in rl
Lee, K.-H., Fischer, I., Liu, A., Guo, Y., Lee, H., Canny, J., and Guadarrama, S · 2020
Cited alongside, same era.
Multi-task reinforcement learning with a planning quasi-metric
Micheli, V., Sinnathamby, K., and Fleuret, F · 2020
Cited alongside, same era.
An inductive bias for distances: Neural nets that respect the triangle inequality
Pitis, S., Chan, H., Jamali, K., and Ba, J · 2020
Cited alongside, same era.
Eysenbach, B., Zhang, T., Salakhutdinov, R., and Levine, S · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Why should I trust you, Bellman? the Bellman error is a poor replacement for value error
Fujimoto, S., Meger, D., Precup, D., Nachum, O., and Gu, S. S · 2022
Later among the works it cites.
Ghasemipour, S. K. S., Gu, S. S., and Nachum, O · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J., and Levine, S · 2022
Later among the works it cites.
Metric residual networks for sample efficient goal-conditioned reinforcement learning
Liu, B., Feng, Y., Liu, Q., and Stone, P · 2022
Later among the works it cites.
Learning dynamics and generalization in reinforcement learning
Lyle, C., Rowland, M., Dabney, W., Kwiatkowska, M., and Gal, Y · 2022
Later among the works it cites.
VIP: Towards universal visual reward and representation via value-implicit pre-training
Ma, Y. J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., and Zhang, A · 2022
Later among the works it cites.
You can’t count on luck: Why decision transformers fail in stochastic environments
Paster, K., McIlraith, S., and Ba, J · 2022
Later among the works it cites.
Improved representation of asymmetrical distances with interval quasimetric embeddings
Wang, T. and Isola, P · 2022
Later among the works it cites.
Denoised MDPs: Learning world models better than the world itself
Wang, T., Du, S. S., Torralba, A., Isola, P., Zhang, A., and Tian, Y · 2022
Later among the works it cites.
Dichotomy of control: Separating what you can control from what you cannot
Yang, M., Schuurmans, D., Abbeel, P., and Nachum, O · 2022
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2023
Closest in time.