Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Deepmind control suite
Original
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Later among the works it cites.
Reward constrained policy optimization
Original
C. Tessler, D. J. Mankowitz, and S. Mannor · 2018
Later among the works it cites.
Reactive reinforcement learning in asynchronous environments
J. B. Travnik, K. W. Mathewson, R. S. Sutton, and P. M. Pilarski · 2018
Later among the works it cites.
Safe exploration and optimization of constrained mdps using gaussian processes
A. Wachi, Y. Sui, Y. Yue, and M. Ono · 2018
Later among the works it cites.
Exponentially weighted imitation learning for batched historical data
Q. Wang, J. Xiong, L. Han, p. sun, H. Liu, and T. Zhang · 2018
Later among the works it cites.
Learn what not to learn: Action elimination with deep reinforcement learning
T. Zahavy, M. Haroush, N. Merlis, D. J. Mankowitz, and S. Mannor · 2018
Later among the works it cites.
Striving for simplicity in off-policy deep reinforcement learning
Original
R. Agarwal, D. Schuurmans, and M. Norouzi · 2019
Later among the works it cites.
ROBEL: RObotics BEnchmarks for Learning with low-cost robots
M. Ahn, H. Zhu, K. Hartikainen, H. Ponte, A. Gupta, S. Levine, and V. Kumar · 2019
Later among the works it cites.
Value constrained model-free continuous control
Original
S. Bohez, A. Abdolmaleki, M. Neunert, J. Buchli, N. Heess, and R. Hadsell · 2019
Later among the works it cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Original
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, O. Sushkov, D. Barker, J. Scholz, M. Denil, N. de Freitas, and Z. Wang · 2019
Later among the works it cites.
Challenges of real-world reinforcement learning
Original
G. Dulac-Arnold, D. J. Mankowitz, and T. Hester · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
Recsim: A configurable simulation platform for recommender systems
Original
E. Ie, C.-w. Hsu, M. Mladenov, V. Jain, S. Narvekar, J. Wang, R. Wu, and C. Boutilier · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Original
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, À. Lapedriza, N. Jones, S. Gu, and R. W. Picard · 2019
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Prediction, consistency, curvature: Representation learning for locally-linear control
Original
N. Levine, Y. Chow, R. Shu, A. Li, M. Ghavamzadeh, and H. Bui · 2019
Later among the works it cites.
Deep Reinforcement Learning for Multi-objective Optimization
K. Li, T. Zhang, and R. Wang · 2019
Later among the works it cites.
Robust reinforcement learning for continuous control with model misspecification
Original
D. J. Mankowitz, N. Levine, R. Jeong, A. Abdolmaleki, J. T. Springenberg, T. A. Mann, T. Hester, and M. A. Riedmiller · 2019
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Original
A. Nagabandi, K. Konoglie, S. Levine, and V. Kumar · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Original
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepezvari, S. Singh, et al · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Original
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Later among the works it cites.
Real-time reinforcement learning
S. Ramstedt and C. Pal · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
A. Ray, J. Achiam, and D. Amodei · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Original
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2019
Later among the works it cites.
Action assembly: Sparse imitation learning for text based games with combinatorial action spaces
Original
C. Tessler, T. Zahavy, D. Cohen, D. J. Mankowitz, and S. Mannor · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
M. Vecerík, O. Sushkov, D. Barker, T. Rothörl, T. Hester, and J. Scholz · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
M. Vecerik, O. Sushkov, D. Barker, T. Rothörl, T. Hester, and J. Scholz · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Original
Y. Wu, G. Tucker, and O. Nachum · 2019
Later among the works it cites.
A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
R. Yang, X. Sun, and K. Narasimhan · 2019
Later among the works it cites.
A distributional view on multi-objective policy optimization
Original
A. Abdolmaleki, S. H. Huang, L. Hasenclever, M. Neunert, H. F. Song, M. Zambelli, M. F. Martins, N. Heess, R. Hadsell, and M. Riedmiller · 2020
Closest in time.
Model-based offline planning
Original
A. Argenson and G. Dulac-Arnold · 2020
Closest in time.
Balancing constraints and rewards with meta-gradient d4pg, 2020
D. A. Calian, D. J. Mankowitz, T. Zahavy, Z. Xu, J. Oh, N. Levine, and T. Mann · 2020
Closest in time.
Acme: A research framework for distributed reinforcement learning
Original
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, et al · 2020
Closest in time.
Morel: Model-based offline reinforcement learning
Original
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Closest in time.
Robust constrained reinforcement learning for continuous control with model misspecification, 2020
D. J. Mankowitz, D. A. Calian, R. Jeong, C. Paduraru, N. Heess, S. Dathathri, M. Riedmiller, and T. Mann · 2020
Closest in time.
Constrained markov decision processes via backward value functions
Original
H. Satija, P. Amortila, and J. Pineau · 2020
Closest in time.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Closest in time.
Responsive safety in reinforcement learning by pid lagrangian methods
Original
A. Stooke, J. Achiam, and P. Abbeel · 2020
Closest in time.
Critic regularized regression
Original
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2020
Closest in time.
Mopo: Model-based offline policy optimization
Original
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Closest in time.