Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine · 2018
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech Marian Czarnecki, Michaël Mathieu, Andrew Joseph Dudzik, Junyoung Chung, Duck Hwan Choi, Richard W. Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander Sasha Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom Le Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy P. Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver · 2019
Later among the works it cites.
Preventing undesirable behavior of intelligent machines
Philip S. Thomas, Bruno Castro da Silva, Andrew G. Barto, Stephen Giguere, Yuriy Brun, and Emma Brunskill · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Original
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Original
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Original
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard · 2019
Later among the works it cites.
Tossingbot: Learning to throw arbitrary objects with residual physics
Original
Andy Zeng, Shuran Song, Johnny Lee, Alberto Rodríguez, and Thomas A. Funkhouser · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Original
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Deep inverse reinforcement learning for sepsis treatment
Chao Yu, Guoqi Ren, and Jiming Liu 0001 · 2019
Later among the works it cites.
Lessons from real-world reinforcement learning in a customer support bot
Original
Nikos Karampatziakis, Sebastian Kochman, Jade Huang, Paul Mineiro, Kathy Osborne, and Weizhu Chen · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Original
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G. Bellemare · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Mo’ states mo’ problems: Emergency stop mechanisms from observation
Original
Samuel K. Ainsworth, Matt Barnes, and Siddhartha S. Srinivasa · 2019
Later among the works it cites.
P3o: Policy-on policy-off policy optimization
Original
Rasool Fakoor, Pratik Chaudhari, and Alexander J. Smola · 2019
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2019
Later among the works it cites.
Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
Kendall Lowrey, Aravind Rajeswaran, Sham Kakade, Emanuel Todorov, and Igor Mordatch · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Original
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, S. Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Original
Anusha Nagabandi, Kurt Konoglie, Sergey Levine, and Vikash Kumar · 2019
Later among the works it cites.
Data efficient reinforcement learning for legged robots
Original
Yuxiang Yang, Ken Caluwaerts, Atil Iscen, Tingnan Zhang, Jie Tan, and Vikas Sindhwani · 2019
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma · 2019
Later among the works it cites.
A Game Theoretic Framework for Model-Based Reinforcement Learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Closest in time.
Mopo: Model-based offline policy optimization
Original
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Closest in time.
Planning and execution using inaccurate models with provable guarantees, 2020
Anirudh Vemula, Yash Oza, J. Andrew Bagnell, and Maxim Likhachev · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning
Original
Justin Fu, Aviral Kumar, Ofir Nachum, G. Tucker, and S. Levine · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
Original
Aviral Kumar, Aurick Zhou, G. Tucker, and S. Levine · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Reinforcement learning via fenchel-rockafellar duality
Original
Ofir Nachum and Bo Dai · 2020
Closest in time.
Exploring model-based planning with policy networks
Original
Tingwu Wang and Jimmy Ba · 2020
Closest in time.