Fetching the paper…
Reading the bibliography…
In this work, we study the use of the Bellman equation as a surrogate objective for value prediction accuracy.
Benchmarking batch deep reinforcement learning algorithms
Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J · 1910
Earlier work this paper cites.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Generalized polynomial approximations in markovian decision processes
Schweitzer, P. J. and Seidmann, A · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Tight performance bounds on greedy policies based on imperfect value functions
Williams, R. J. and Baird, L · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Singh, S. P. and Yee, R. C · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G · 1996
Earlier work this paper cites.
The loss from imperfect value functions in expectation-based and minimax-based tasks
Heger, M · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Performance bounds in l_p-norm for approximate value iteration
Munos, R · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A. M., Munos, R., and Szepesvári, C · 2010
Earlier work this paper cites.
Model selection in reinforcement learning
Farahmand, A.-m. and Szepesvári, C · 2011
Earlier work this paper cites.
The fixed points of off-policy td
Kolter, J. Z · 2011
Cited alongside, same era.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
Deterministic bellman residual minimization
Saleh, E. and Jiang, N · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z. D., and Blundell, C · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Is the bellman residual a bad proxy?
Geist, M., Piot, B., and Pietquin, O · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D., and Kakade, S. M · 2020
Later among the works it cites.
Deep residual reinforcement learning
Zhang, S., Boehmer, W., and Whiteson, S · 2020
Later among the works it cites.
Borrowing from the future: An attempt to address double sampling
Zhu, Y. and Ying, L · 2020
Later among the works it cites.
Logistic q-learning
Bas-Serrano, J., Curi, S., Krause, A., and Neu, G · 2021
Later among the works it cites.
Chen, L., Scherrer, B., and Bartlett, P. L · 2021
Later among the works it cites.
Benchmarks for deep off-policy evaluation
Fu, J., Norouzi, M., Nachum, O., Tucker, G., Wang, Z., Novikov, A., Yang, M., Zhang, M. R., Chen, Y., Kumar, A., Paduraru, C., Levine, S., and Paine, T · 2021
Later among the works it cites.
A deep reinforcement learning approach to marginalized importance sampling with the successor representation
Fujimoto, S., Meger, D., and Precup, D · 2021
Later among the works it cites.
Model selection for offline reinforcement learning: Practical considerations for healthcare settings
Tang, S. and Wiens, J · 2021
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Voloshin, C., Le, H. M., Jiang, N., and Yue, Y · 2021
Later among the works it cites.
On the sample complexity of batch reinforcement learning with policy-induced data
Xiao, C., Lee, I., Dai, B., Schuurmans, D., and Szepesvari, C · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Yang, M. and Nachum, O · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Zanette, A · 2021
Later among the works it cites.
A generalized projected bellman error for off-policy value estimation in reinforcement learning
Patterson, A., White, A., and White, M · 2022
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2062
Closest in time.