Fetching the paper…
Reading the bibliography…
The pre-train and fine-tune paradigm in machine learning has had dramatic success in a wide range of domains because the use of existing data or pre-trained models on the internet enables quick and easy learning of new tasks.
A. Y. Ng, S. Russell, et al. , “Algorithms for inverse reinforcement learning.” in Icml , vol. 1, 2000, p. 2
2000
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning , 2004, p. 1
2004
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al. , “Maximum entropy inverse reinforcement learning.” in Aaai , vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
V. Kumar, E. Todorov, and S. Levine, “Optimal control with learned local models: Application to dexterous manipulation,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2016, pp. 378–383
2016
Earlier work this paper cites.
J. Ho, J. Gupta, and S. Ermon, “Model-free imitation learning with policy optimization,” in International conference on machine learning . PMLR, 2016, pp. 2760–2769
2016
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning . PMLR, 2016, pp. 49–58
2016
Earlier work this paper cites.
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman, “Deep predictive policy training using reinforcement learning,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2017, pp. 2351–2358
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY, USA: Curran Associates Inc., 2018, p. 8547–8556
2018
Earlier work this paper cites.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain, “Time-contrastive networks: Self-supervised learning from video,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 1134–1141
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “Mopo: Model-based offline policy optimization,” Advances in Neural Information Processing Systems , vol. 33, pp. 14 129–14 142, 2020
2020
Cited alongside, same era.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 1179–1191, 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Lyu, X. Ma, X. Li, and Z. Lu, “Mildly conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 1711–1724, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
K. Ploeger, M. Lutter, and J. Peters, “High acceleration reinforcement learning for real-world juggling with binary rewards,” in Conference on Robot Learning . PMLR, 2021, pp. 642–653
2021
Cited alongside, same era.
A. Gupta, J. Yu, T. Z. Zhao, V. Kumar, A. Rovinsky, K. Xu, T. Devlin, and S. Levine, “Reset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 6664–6671
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Wu, H. Wu, Z. Qiu, J. Wang, and M. Long, “Supported policy optimization for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 31 278–31 291, 2022
2022
Later among the works it cites.
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” in Conference on Robot Learning . PMLR, 2022, pp. 1702–1712
2022
Later among the works it cites.
M. S. Mark, A. Ghadirzadeh, X. Chen, and C. Finn, “Fine-tuning offline policies with optimistic action selection,” in Deep Reinforcement Learning Workshop NeurIPS 2022 , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
H. R. Walke, J. H. Yang, A. Yu, A. Kumar, J. Orbik, A. Singh, and S. Levine, “Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning,” in Conference on Robot Learning . PMLR, 2023, pp. 1652–1662
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
K. Xu, Z. Hu, R. Doshi, A. Rovinsky, V. Kumar, A. Gupta, and S. Levine, “Dexterous manipulation from images: Autonomous real-world rl via substep guidance,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 5938–5945
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.