Fetching the paper…
Reading the bibliography…
The primacy bias in model-free reinforcement learning (MFRL), which refers to the agent's tendency to overfit early data and lose the ability to learn from new data, can significantly decrease the performance of MFRL algorithms.
Philip H Marshall and Pamela R Werder, ‘The effects of the elimination of rehearsal on primacy and recency’, Journal of Verbal Learning and Verbal Behavior
1972
Earlier work this paper cites.
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel, ‘Backpropagation applied to handwritten zip code recognition’, Neural computation
1989
Earlier work this paper cites.
Richard S Sutton, ‘Dyna, an integrated architecture for learning, planning, and reacting’, ACM Sigart Bulletin
1991
Earlier work this paper cites.
Sebastian Thrun, ‘Lifelong learning algorithms’, in Learning to learn
1998
Earlier work this paper cites.
Nicolas Schweighofer and Kenji Doya, ‘Meta-learning in reinforcement learning’, Neural Networks
2003
Earlier work this paper cites.
Emanuel Todorov, Tom Erez, and Yuval Tassa, ‘Mujoco: A physics engine for model-based control’, in 2012 IEEE/RSJ international conference on intelligent robots and systems
2012
Earlier work this paper cites.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, ‘Dropout: a simple way to prevent neural networks from overfitting’, The journal of machine learning research
2014
Earlier work this paper cites.
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis, ‘The dependence of effective planning horizon on model accuracy’, in Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems
2015
Earlier work this paper cites.
Franziska Meier and Stefan Schaal, ‘Drifting gaussian processes with varying neighborhood sizes for online model learning’, in International Conference on Robotics and Automation (ICRA)
2016
Earlier work this paper cites.
Amir-massoud Farahmand, Andre Barreto, and Daniel Nikovski, ‘Value-aware loss function for model-based reinforcement learning’, in Artificial Intelligence and Statistics
2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, ‘Attention is all you need’, Advances in neural information processing systems
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine, ‘Deep reinforcement learning in a handful of trials using probabilistic dynamics models’, Advances in neural information processing systems
2018
Earlier work this paper cites.
Scott Fujimoto, Herke Hoof, and David Meger, ‘Addressing function approximation error in actor-critic methods’, in International conference on machine learning
2018
Earlier work this paper cites.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine, ‘Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor’, in International conference on machine learning
2018
Earlier work this paper cites.
Anusha Nagabandi, Ignasi Clavera, Simin Liu, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn, ‘Learning to adapt in dynamic, real-world environments through meta-reinforcement learning’, in International Conference on Learning Representations
2018
Earlier work this paper cites.
Richard S Sutton and Andrew G Barto, Reinforcement learning: An introduction
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Scott Fujimoto, David Meger, and Doina Precup, ‘Off-policy deep reinforcement learning without exploration’, in International conference on machine learning
2019
Cited alongside, same era.
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka, ‘Dropout q-functions for doubly efficient reinforcement learning’, in International Conference on Learning Representations
2021
Later among the works it cites.
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A Al Sallab, Senthil Yogamani, and Patrick Pérez, ‘Deep reinforcement learning for autonomous driving: A survey’, IEEE Transactions on Intelligent Transportation Systems
2021
Later among the works it cites.
2021
Later among the works it cites.
Hang Lai, Jian Shen, Weinan Zhang, Yimin Huang, Xing Zhang, Ruiming Tang, Yong Yu, and Zhenguo Li, ‘On effective scheduling of model-based reinforcement learning’, Advances in Neural Information Processing Systems
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine, ‘When to trust your model: Model-based policy optimization’, Advances in neural information processing systems
2019
Cited alongside, same era.
Łukasz Kaiser, Mohammad Babaeizadeh, Piotr Miłos, Błażej Osiński, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al., ‘Model based reinforcement learning for atari’, in International Conference on Learning Representations
2019
Cited alongside, same era.
Hado P Van Hasselt, Matteo Hessel, and John Aslanides, ‘When to use parametric models in reinforcement learning?’, Advances in Neural Information Processing Systems
2019
Cited alongside, same era.
Xinyue Chen, Che Wang, Zijian Zhou, and Keith W Ross, ‘Randomized ensembled double q-learning: Learning fast without a model’, in International Conference on Learning Representations
2020
Cited alongside, same era.
William Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio, Hugo Larochelle, Mark Rowland, and Will Dabney, ‘Revisiting fundamentals of experience replay’, in International Conference on Machine Learning
2020
Cited alongside, same era.
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba, ‘Mastering atari with discrete world models’, in International Conference on Learning Representations
2020
Cited alongside, same era.
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine, ‘Conservative q-learning for offline reinforcement learning’, Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
Kimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee, and Jinwoo Shin, ‘Context-aware dynamics model for generalization in model-based reinforcement learning’, in International Conference on Machine Learning
2020
Cited alongside, same era.
Clare Lyle, Mark Rowland, and Will Dabney, ‘Understanding and preventing capacity loss in reinforcement learning’, in International Conference on Learning Representations
2021
Later among the works it cites.
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig, ‘Safe learning in robotics: From learning-based control to safe reinforcement learning’, Annual Review of Control, Robotics, and Autonomous Systems
2022
Later among the works it cites.
Edoardo Cetin, Philip J Ball, Stephen Roberts, and Oya Celiktutan, ‘Stabilizing off-policy deep reinforcement learning from pixels’, in International Conference on Machine Learning
2022
Later among the works it cites.
Pierluca D’Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G Bellemare, and Aaron Courville, ‘Sample-efficient reinforcement learning by breaking the replay ratio barrier’, in Deep Reinforcement Learning Workshop NeurIPS 2022
2022
Later among the works it cites.
Tianying Ji, Yu Luo, Fuchun Sun, Mingxuan Jing, Fengxiang He, and Wenbing Huang, ‘When to update your model: Constrained model-based reinforcement learning’, Advances in Neural Information Processing Systems
2022
Later among the works it cites.
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup, ‘Towards continual reinforcement learning: A review and perspectives’, Journal of Artificial Intelligence Research
2022
Later among the works it cites.
Qiyang Li, Aviral Kumar, Ilya Kostrikov, and Sergey Levine, ‘Efficient deep reinforcement learning requires regulating overfitting’, in The Eleventh International Conference on Learning Representations
2022
Later among the works it cites.
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville, ‘The primacy bias in deep reinforcement learning’, in International Conference on Machine Learning
2022
Later among the works it cites.
2022
Later among the works it cites.
Ruijie Zheng, Xiyao Wang, Huazhe Xu, and Furong Huang, ‘Is model ensemble necessary? model-based rl via a single model with lipschitz regularized value function’, in The Eleventh International Conference on Learning Representations
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
Xiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, and Furong Huang, ‘Live in the moment: Learning dynamics model adapted to evolving policy’, in International Conference on Machine Learning
2023
Closest in time.