Fetching the paper…
Reading the bibliography…
Substantial advancements to model-based reinforcement learning algorithms have been impeded by the model-bias induced by the collected data, which generally hurts performance.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
C. G. Atkeson, “Using local trajectory optimizers to speed up global optimization in dynamic programming,” in Advances in neural information processing systems , 1994, pp. 663–670
1994
Earlier work this paper cites.
Y. Tassa, T. Erez, and W. D. Smart, “Receding horizon differential dynamic programming,” in Advances in neural information processing systems , 2008, pp. 1465–1472
2008
Earlier work this paper cites.
M. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on machine learning (ICML-11) , 2011, pp. 465–472
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning , 2013, pp. 1–9
2013
Earlier work this paper cites.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in Advances in Neural Information Processing Systems , 2014, pp. 1071–1079
2014
Earlier work this paper cites.
R. R. Ma and A. M. Dollar, “An underactuated hand for efficient finger-gaiting-based dexterous manipulation,” in 2014 IEEE International Conference on Robotics and Biomimetics (ROBIO 2014) , 2014, pp. 2214–2219
2014
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning , 2015, pp. 1889–1897
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
B. Calli, A. Walsman, A. Singh, S. Srinivasa, P. Abbeel, and A. M. Dollar, “Benchmarking in manipulation research: Using the yale-cmu-berkeley object and model set,” IEEE Robotics & Automation Magazine , vol. 22, no. 3, pp. 36–52, 2015
2015
Earlier work this paper cites.
V. Kumar, E. Todorov, and S. Levine, “Optimal control with learned local models: Application to dexterous manipulation,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2016, pp. 378–383
2016
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334–1373, 2016
2016
Earlier work this paper cites.
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou, “Aggressive driving with model predictive path integral control,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2016, pp. 1433–1440
2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” 2016
2016
Cited alongside, same era.
J. Wang and E. Olson, “Apriltag 2: Efficient and robust fiducial detection,” in 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2016, pp. 4193–4198
2016
Cited alongside, same era.
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou, “Information theoretic mpc for model-based reinforcement learning,” in 2017 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2017, pp. 1714–1721
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Later among the works it cites.
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 7559–7566
2018
Later among the works it cites.
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” in 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 2017, pp. 3389–3396
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Advances in Neural Information Processing Systems , 2018, pp. 4754–4765
2018
Cited alongside, same era.
B. Amos, I. Jimenez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable mpc for end-to-end planning and control,” in Advances in Neural Information Processing Systems , 2018, pp. 8289–8300
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Haarnoja, V. Pong, A. Zhou, M. Dalal, P. Abbeel, and S. Levine, “Composable deep reinforcement learning for robotic manipulation,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 6244–6251
2018
Cited alongside, same era.
2018
Later among the works it cites.
G. Chalvatzaki, X. S. Papageorgiou, P. Maragos, and C. S. Tzafestas, “Learn to adapt to human walking: A model-based reinforcement learning approach for a robotic assistant rollator,” IEEE Robotics and Automation Letters , vol. 4, no. 4, pp. 3774–3781, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Janner, J. Fu, M. Zhang, and S. Levine, “When to trust your model: Model-based policy optimization,” in Advances in Neural Information Processing Systems , 2019, pp. 12 519–12 530
2019
Later among the works it cites.
2019
Later among the works it cites.
A. S. Morgan, K. Hang, and A. M. Dollar, “Object-agnostic dexterous manipulation of partially constrained trajectories,” IEEE Robotics and Automation Letters , vol. 5, no. 4, pp. 5494–5501, 2020
2020
Later among the works it cites.
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar, “Deep dynamics models for learning dexterous manipulation,” in Conference on Robot Learning , 2020, pp. 1101–1112
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Bhardwaj, A. Handa, D. Fox, and B. Boots, “Information theoretic model predictive q-learning,” in Learning for Dynamics and Control , 2020, pp. 840–850
2020
Later among the works it cites.
A. S. Morgan, W. G. Bircher, and A. M. Dollar, “Towards generalized manipulation learning through grasp mechanics-based features and self-supervision,” IEEE Transactions on Robotics , pp. 1–17, 2021
2021
Closest in time.