Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) solves sequential decision-making problems via a trial-and-error process interacting with the environment.
Dynamic programming and stochastic control processes
R. Bellman · 1958
Earlier work this paper cites.
Equilibrium in a stochastic n-person game
A. M. Fink · 1964
Earlier work this paper cites.
Linear optimal control systems , volume 1072
H. Kwakernaak and R. Sivan · 1969
Earlier work this paper cites.
Markov processes over denumerable products of spaces, describing large systems of automata
L. N. Vaserstein · 1969
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Bayesian methods for adaptive models
D. J. C. Mackay · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
A. W. Moore and C. G. Atkeson · 1993
Earlier work this paper cites.
Temporal difference learning and td-gammon
G. Tesauro · 1995
Earlier work this paper cites.
On-line policy improvement using monte-carlo search
G. Tesauro and G. R. Galperin · 1996
Earlier work this paper cites.
State abstraction in MAXQ hierarchical reinforcement learning
T. G. Dietterich · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. J. Kearns and S. P. Singh · 2002
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
Gaussian processes for machine learning
M. W. Seeger · 2004
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
E. Todorov and W. Li · 2005
Earlier work this paper cites.
A framework for reinforcement learning and planning
T. M. Moerland, J. Broekens, and C. M. Jonker · 2006
Earlier work this paper cites.
Model-based reinforcement learning: A survey
T. M. Moerland, J. Broekens, and C. M. Jonker · 2006
Earlier work this paper cites.
Policy gradient methods for robotics
J. Peters and S. Schaal · 2006
Earlier work this paper cites.
Computing "elo ratings" of move patterns in the game of go
R. Coulom · 2007
Earlier work this paper cites.
Apprenticeship learning using linear programming
U. Syed, M. Bowling, and R. E. Schapire · 2008
Earlier work this paper cites.
Continuous upper confidence trees
A. Couëtoux, J. Hoock, N. Sokolovska, O. Teytaud, and N. Bonnard · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. P. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. Browne, E. J. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. P. Liebana, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Partially observable Markov decision processes
M. T. Spaan · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Y. Tassa, T. Erez, and E. Todorov · 2012
Earlier work this paper cites.
The cross-entropy method for optimization
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer · 2013
Earlier work this paper cites.
Model Predictive Control
E. F. Camacho and C. B. Alba · 2013
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
M. Hausknecht and P. Stone · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. P. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Abstraction selection in model-based reinforcement learning
N. Jiang, A. Kulesza, and S. Singh · 2015
Earlier work this paper cites.
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Optimizing the cvar via sampling
A. Tamar, Y. Glassner, and S. Mannor · 2015
Earlier work this paper cites.
Improving multi-step prediction of learned time series models
A. Venkatraman, M. Hebert, and J. A. Bagnell · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. T. Springenberg, J. Boedecker, and M. A. Riedmiller · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. A. Theodorou · 2015
Earlier work this paper cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Improving pilco with bayesian neural network dynamics models
Y. Gal, R. McAllister, and C. E. Rasmussen · 2016
Earlier work this paper cites.
The CMA evolution strategy: A tutorial
N. Hansen · 2016
Earlier work this paper cites.
Opponent modeling in deep reinforcement learning
H. He, J. Boyd-Graber, K. Kwok, and H. Daumé III · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Value iteration networks
A. Tamar, S. Levine, P. Abbeel, Y. Wu, and G. Thomas · 2016
Earlier work this paper cites.
Derivative-free optimization via classification
Y. Yu, H. Qian, and Y. Hu · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Cited alongside, same era.
Value-aware loss function for model-based reinforcement learning
A.-m. Farahmand, A. Barreto, and D. Nikovski · 2017
Cited alongside, same era.
A benchmark environment motivated by industrial control problems
D. Hein, S. Depeweg, M. Tokic, S. Udluft, A. Hentschel, T. A. Runkler, and V. Sterzing · 2017
Cited alongside, same era.
Sequential classification-based optimization for direct policy search
Y. Hu, H. Qian, and Y. Yu · 2017
Cited alongside, same era.
Policy-aware model learning for policy gradient methods
R. Abachi · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Y. Bai and C. Jin · 2020
Later among the works it cites.
Model-predictive control via cross-entropy and gradient-based optimization
H. Bharadhwaj, K. Xie, and F. Shkurti · 2020
Later among the works it cites.
Bail: Best-action imitation learning for batch deep reinforcement learning
X. Chen, Z. Zhou, Z. Wang, C. Wang, Y. Wu, and K. Ross · 2020
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
I. Clavera, Y. Fu, and P. Abbeel · 2020
Later among the works it cites.
Intelligent trainer for dyna-style model-based deep reinforcement learning
L. Dong, Y. Li, X. Zhou, Y. Wen, and K. Guan · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagination-augmented agents for deep reinforcement learning
S. Racanière, T. Weber, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, R. Pascanu, P. W. Battaglia, D. Hassabis, D. Silver, and D. Wierstra · 2017
Cited alongside, same era.
Sim-to-real robot learning from pixels with progressive nets
A. A. Rusu, M. Večerík, T. Rothörl, N. Heess, R. Pascanu, and R. Hadsell · 2017
Cited alongside, same era.
Preparing for the unknown: Learning a universal policy with online system identification
W. Yu, J. Tan, C. K. Liu, and G. Turk · 2017
Cited alongside, same era.
Lipschitz continuity in model-based reinforcement learning
K. Asadi, D. Misra, and M. L. Littman · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
Using control synthesis to generate corner cases: A case study on autonomous driving
G. Chou, Y. E. Sahin, L. Yang, K. J. Rutledge, P. Nilsson, and N. Ozay · 2018
Cited alongside, same era.
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2020
Later among the works it cites.
Learning-based model predictive control: Toward safe learning in control
L. Hewing, K. P. Wabersich, M. Menner, and M. N. Zeilinger · 2020
Later among the works it cites.
Notes on Rmax exploration, 2020
N. Jiang · 2020
Later among the works it cites.
Offline imitation learning with a misspecified simulator
S. Jiang, J.-C. Pang, and Y. Yu · 2020
Later among the works it cites.
MoRel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Communication in multi-agent reinforcement learning: Intention sharing
W. Kim, J. Park, and Y. Sung · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
N. Lambert, B. Amos, O. Yadan, and R. Calandra · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
FinRL: A deep reinforcement learning library for automated stock trading in quantitative finance
X.-Y. Liu, H. Yang, Q. Chen, R. Zhang, L. Yang, B. Xiao, and C. D. Wang · 2020
Later among the works it cites.
Monte carlo gradient estimation in machine learning
S. Mohamed, M. Rosca, M. Figurnov, and A. Mnih · 2020
Later among the works it cites.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
S. Nair and C. Finn · 2020
Later among the works it cites.
Trust the model when it is confident: Masked model-based actor-critic
F. Pan, J. He, D. Tu, and Q. He · 2020
Later among the works it cites.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
S. Pitis, H. Chan, S. Zhao, B. Stadie, and J. Ba · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
A. Rajeswaran, I. Mordatch, and V. Kumar · 2020
Later among the works it cites.
Model-based policy optimization with unsupervised model adaptation
J. Shen, H. Zhao, W. Zhang, and Y. Yu · 2020
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2020
Later among the works it cites.
On computation and generalization of generative adversarial imitation learning
Y. Wang, T. Liu, Z. Yang, X. Li, Z. Wang, and T. Zhao · 2020
Later among the works it cites.
Model primitives for hierarchical lifelong reinforcement learning
B. Wu, J. Gupta, and M. Kochenderfer · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
T. Xu, Z. Li, and Y. Yu · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
On the model-based stochastic value gradient for continuous reinforcement learning
B. Amos, S. Stanton, D. Yarats, and A. G. Wilson · 2021
Later among the works it cites.
Model-based meta-reinforcement learning for flight with suspended payloads
S. Belkhale, R. Li, G. Kahn, R. McAllister, R. Calandra, and S. Levine · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. S. Chatterji, A. S. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. D. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. S. Krass, R. Krishna, R. Kuditipudi, and et al · 2021
Later among the works it cites.
Mismatched no more: Joint model-policy optimization for model-based RL
B. Eysenbach, A. Khazatsky, S. Levine, and R. Salakhutdinov · 2021
Later among the works it cites.
Mastering Atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2021
Later among the works it cites.
Simgan: Hybrid simulator identification for domain adaptation via adversarial reinforcement learning
Y. Jiang, T. Zhang, D. Ho, Y. Bai, C. K. Liu, S. Levine, and J. Tan · 2021
Later among the works it cites.
On effective scheduling of model-based reinforcement learning
H. Lai, J. Shen, W. Zhang, Y. Huang, X. Zhang, R. Tang, Y. Yu, and Z. Li · 2021
Later among the works it cites.
Tesseract: Tensorised actors for multi-agent reinforcement learning
A. Mahajan, M. Samvelyan, L. Mao, V. Makoviychuk, A. Garg, J. Kossaifi, S. Whiteson, Y. Zhu, and A. Anandkumar · 2021
Later among the works it cites.
Neorl: A near real-world benchmark for offline reinforcement learning
R. Qin, S. Gao, X. Zhang, Z. Xu, S. Huang, Z. Li, W. Zhang, and Y. Yu · 2021
Later among the works it cites.
Partially observable environment estimation with uplift inference for reinforcement learning based recommendation, 2021
W. Shang, Q. Li, Z. Qin, Y. Yu, Y. Meng, and J. Ye · 2021
Later among the works it cites.
Robustness and sample complexity of model-based MARL for general-sum markov games
J. Subramanian, A. Sinha, and A. Mahajan · 2021
Later among the works it cites.
Corner case generation and analysis for safety assessment of autonomous vehicles
H. Sun, S. Feng, X. Yan, and H. X. Liu · 2021
Later among the works it cites.
Value gradient weighted model-based reinforcement learning
C. A. Voelcker, V. Liao, A. Garg, and A.-m. Farahmand · 2021
Later among the works it cites.
Offline reinforcement learning with reverse model-based imagination
J. Wang, W. Li, H. Jiang, G. Zhu, S. Li, and C. Zhang · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
M. Yang and O. Nachum · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Later among the works it cites.
Smarts: An open-source scalable multi-agent rl training school for autonomous driving
M. Zhou, J. Luo, J. Villella, Y. Yang, D. Rusu, J. Miao, W. Zhang, M. Alban, I. FADAKAR, Z. Chen, et al · 2021
Later among the works it cites.
Mapgo: Model-assisted policy optimization for goal-oriented tasks
M. Zhu, M. Liu, J. Shen, Z. Zhang, S. Chen, W. Zhang, D. Ye, Y. Yu, Q. Fu, and W. Yang · 2021
Later among the works it cites.
Adversarial counterfactual environment model learning
X.-H. Chen, Y. Yu, Z.-M. Zhu, Z. Yu, Z. Chen, C. Wang, Y. Wu, H. Wu, R.-J. Qin, R. Ding, and F. Huang · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning, 2022
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de las Casas, C. Donner, L. Fritz, C. Galperti, A. Huber, J. Keeling, M. Tsimpoukelli, J. Kay, A. Merle, J.-M. Moret, S. Noury, F. Pesamosca, D. Pfau, O. Sauter, C. Sommariva, S. Coda, B. Duval, A. Fasoli, P. Kohli, K. Kavukcuoglu, D. Hassabis, and M. Riedmiller · 2022
Closest in time.
Hybrid value estimation for off-policy evaluation and offline reinforcement learning
X.-K. Jin, X.-H. Liu, S. Jiang, and Y. Yu · 2022
Closest in time.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Closest in time.
Adapt to environment sudden changes by learning a context sensitive policy
F.-M. Luo, S. Jiang, Y. Yu, Z. Zhang, and Y.-F. Zhang · 2022
Closest in time.
Learning robust perceptive locomotion for quadrupedal robots in the wild
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2022
Closest in time.
S. E. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Closest in time.
Model-based multi-agent reinforcement learning: Recent progress and prospects
X. Wang, Z. Zhang, and W. Zhang · 2022
Closest in time.