Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is typically concerned with estimating stationary policies or single-step models, leveraging the Markov property to factorize problems in time.
Kumar, A., Peng, X. B., and Levine, S · 1912
Earlier work this paper cites.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Speech understanding systems: Summary of results of the five-year research effort at Carnegie Mellon University, 1977
Reddy, R · 1977
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bakker, B · 2002
Earlier work this paper cites.
Kumar, S., Parker, J., and Naderian, P · 2004
Earlier work this paper cites.
Reinforcement learning by value gradients
Fairbank, M · 2008
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
Silver, D., Sutton, R. S., and Müller, M · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Approximate model-assisted neural fitted Q-iteration
Lampe, T. and Riedmiller, M · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Control of memory, active perception, and action in Minecraft
Oh, J., Chockalingam, V., Lee, H., et al · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S · 2017
Earlier work this paper cites.
DeepLoco: Dynamic locomotion skills using hierarchical deep reinforcement learning
Peng, X. B., Berseth, G., Yin, K., and Van De Panne, M · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
Co-Reyes, J., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
γ \gamma -models: Generative temporal difference learning for infinite-horizon prediction
Janner, M., Mordatch, I., and Levine, S · 2020
Later among the works it cites.
minGPT: A minimal pytorch re-implementation of the openai gpt training, 2020
Karpathy, A · 2020
Later among the works it cites.
MOReL: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Later among the works it cites.
If beam search is the answer, what was the question?
Meister, C., Cotterell, R., and Vieira, T · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., S. Fearing, R., and Levine, S · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Insertion-based Decoding with Automatically Inferred Generation Order
Gu, J., Liu, Q., and Cho, K · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Cited alongside, same era.
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2020
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B., Stanton, S., Yarats, D., and Wilson, A. G · 2021
Closest in time.
Model-based offline planning
Argenson, A. and Dulac-Arnold, G · 2021
Closest in time.
Decision Transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Closest in time.
Offline reinforcement learning with pseudometric learning
Dadashi, R., Rezaeifar, S., Vieillard, N., Hussenot, L., Pietquin, O., and Geist, M · 2021
Closest in time.
EMaQ: Expected-max Q-learning operator for simple yet effective offline and online RL
Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S · 2021
Closest in time.
Learning to reach goals via iterated supervised learning
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C. M., Eysenbach, B., and Levine, S · 2021
Closest in time.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., and Wang, Z · 2021
Closest in time.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Closest in time.
Efficient transformers in reinforcement learning using actor-learner distillation
Parisotto, E. and Salakhutdinov, R · 2021
Closest in time.
Planning from pixels using inverse dynamics models
Paster, K., McIlraith, S. A., and Ba, J · 2021
Closest in time.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
Yin, M., Bai, Y., and Wang, Y.-X · 2021
Closest in time.
Autoregressive dynamics models for offline policy evaluation and optimization
Zhang, M. R., Paine, T., Nachum, O., Paduraru, C., Tucker, G., ziyu wang, and Norouzi, M · 2021
Closest in time.