Fetching the paper…
Reading the bibliography…
Offline inverse reinforcement learning (Offline IRL) aims to recover the structure of rewards and environment dynamics that underlie observed actions in a fixed, finite set of demonstrations from an expert agent.
A. Y. Ng, S. Russell et al. , “Algorithms for inverse reinforcement learning.” in Icml , vol. 1, 2000, p. 2
2000
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning , 2004, p. 1
2004
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey et al. , “Maximum entropy inverse reinforcement learning.” in AAAI , vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
E. Klein, M. Geist, and O. Pietquin, “Batch, off-policy and model-free apprenticeship learning,” in European Workshop on Reinforcement Learning . Springer, 2011, pp. 285–296
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
E. Klein, M. Geist, B. Piot, and O. Pietquin, “Inverse reinforcement learning through structured classification,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Earlier work this paper cites.
B. D. Ziebart, J. A. Bagnell, and A. K. Dey, “The principle of maximum causal entropy for estimating interacting processes,” IEEE Transactions on Information Theory , vol. 59, no. 4, pp. 1966–1980, 2013
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Herman, T. Gindele, J. Wagner, F. Schmitt, and W. Burgard, “Inverse reinforcement learning with simultaneous estimation of rewards and dynamics,” in Artificial Intelligence and Statistics . PMLR, 2016, pp. 102–110
2016
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning . PMLR, 2016, pp. 49–58
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Zhou, M. Bloem, and N. Bambos, “Infinite time horizon maximum causal entropy inverse reinforcement learning,” IEEE Transactions on Automatic Control , vol. 63, no. 9, pp. 2787–2802, 2017
2017
Earlier work this paper cites.
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in International conference on machine learning . PMLR, 2017, pp. 1352–1361
2017
Earlier work this paper cites.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel et al. , “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science , vol. 362, no. 6419, pp. 1140–1144, 2018
2018
Earlier work this paper cites.
J. Fu, A. Korattikara, S. Levine, and S. Guadarrama, “From language to goals: Inverse reinforcement learning for vision-based instruction following,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning . PMLR, 2018, pp. 1861–1870
2018
Earlier work this paper cites.
J. Bhandari, D. Russo, and R. Singal, “A finite time analysis of temporal difference learning with linear function approximation,” in Conference on learning theory . PMLR, 2018, pp. 1691–1692
2018
Earlier work this paper cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” Nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Chen, Y. Wang, T. Liu, Z. Yang, X. Li, Z. Wang, and T. Zhao, “On computation and generalization of generative adversarial imitation learning,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Agarwal, N. Jiang, S. M. Kakade, and W. Sun, “Reinforcement learning: Theory and algorithms,” CS Dept., UW Seattle, Seattle, WA, USA, Tech. Rep , pp. 10–4, 2019
2019
Cited alongside, same era.
S. Liu, K. C. See, K. Y. Ngiam, L. A. Celi, X. Sun, M. Feng et al. , “Reinforcement learning for clinical decision support in critical care: comprehensive review,” Journal of medical Internet research , vol. 22, no. 7, p. e18477, 2020
2020
Cited alongside, same era.
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn, “Offline reinforcement learning from images with latent space models,” in Learning for Dynamics and Control . PMLR, 2021, pp. 1154–1168
2021
Later among the works it cites.
C. Lu, P. Ball, J. Parker-Holder, M. Osborne, and S. J. Roberts, “Revisiting design choices in offline model based reinforcement learning,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi, “Fast global convergence of natural policy gradient methods with entropy regularization,” Operations Research , 2021
2021
Later among the works it cites.
M. Uehara and W. Sun, “Pessimistic model-based offline reinforcement learning under partial coverage,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 1179–1191, 2020
2020
Cited alongside, same era.
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “Morel: Model-based offline reinforcement learning,” Advances in neural information processing systems , vol. 33, pp. 21 810–21 823, 2020
2020
Cited alongside, same era.
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “Mopo: Model-based offline policy optimization,” Advances in Neural Information Processing Systems , vol. 33, pp. 14 129–14 142, 2020
2020
Cited alongside, same era.
Z. Wu, L. Sun, W. Zhan, C. Yang, and M. Tomizuka, “Efficient sampling-based maximum entropy inverse reinforcement learning with application to autonomous driving,” IEEE Robotics and Automation Letters , vol. 5, no. 4, pp. 5355–5362, 2020
2020
Cited alongside, same era.
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill, “Provably good batch off-policy reinforcement learning without great exploration,” Advances in neural information processing systems , vol. 33, pp. 1264–1274, 2020
2020
Cited alongside, same era.
Y. F. Wu, W. Zhang, P. Xu, and Q. Gu, “A finite-time analysis of two time-scale actor-critic methods,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 617–17 628, 2020
2020
Cited alongside, same era.
C. Jin, P. Netrapalli, and M. Jordan, “What is local optimality in nonconvex-nonconcave minimax optimization?” in International conference on machine learning . PMLR, 2020, pp. 4880–4889
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z. Guan, T. Xu, and Y. Liang, “When will generative adversarial imitation learning algorithms attain global convergence,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 1117–1125
2021
Later among the works it cites.
H. Cao, S. Cohen, and L. Szpruch, “Identifiability in inverse reinforcement learning,” Advances in Neural Information Processing Systems , vol. 34, pp. 12 362–12 373, 2021
2021
Later among the works it cites.
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn, “Visual adversarial imitation learning using variational models,” Advances in Neural Information Processing Systems , vol. 34, pp. 3016–3028, 2021
2021
Later among the works it cites.
S. Levine, “Understanding the world through action,” in Conference on Robot Learning . PMLR, 2022, pp. 1752–1757
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Wei, A. Garcia, A. McDonald, G. Markkula, J. Engström, I. Supeene, and M. O’Kelly, “World model learning from demonstrations with active inference: Application to driving behavior,” in Forthcoming, 3rd International Workshop on Active Inference, Grenoble, France , 2022
2022
Later among the works it cites.
S. Zeng, C. Li, A. Garcia, and M. Hong, “Maximum-likelihood inverse reinforcement learning with finite-time guarantees,” Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
G. Tennenholtz and S. Mannor, “Uncertainty estimation using riemannian model dynamics for offline reinforcement learning,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. [Online]. Available: https://openreview.net/forum?id=pGLFkjgVvVe
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Shani, T. Zahavy, and S. Mannor, “Online apprenticeship learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 8, 2022, pp. 8240–8248
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Liu and M. Zhu, “Distributed inverse constrained reinforcement learning for multi-agent systems,” Advances in Neural Information Processing Systems , vol. 35, pp. 33 444–33 456, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
C.-A. Cheng, T. Xie, N. Jiang, and A. Agarwal, “Adversarially trained actor critic for offline reinforcement learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 3852–3878
2022
Later among the works it cites.
S. Yue, G. Wang, W. Shao, Z. Zhang, S. Lin, J. Ren, and J. Zhang, “CLARE: Conservative model-based reward learning for offline inverse reinforcement learning,” in International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=5aT4ganOd98
2023
Closest in time.
2023
Closest in time.
M. Hong, H.-T. Wai, Z. Wang, and Z. Yang, “A two-timescale stochastic algorithm framework for bilevel optimization: Complexity analysis and application to actor-critic,” SIAM Journal on Optimization , vol. 33, no. 1, pp. 147–180, 2023
2023
Closest in time.