Fetching the paper…
Reading the bibliography…
We study the problem of Inverse Reinforcement Learning (IRL) with an average-reward criterion.
On-line optimization of simulated markovian processes
G Ch Pflug · 1990
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
Vivek S Borkar and Vijaymohan R Konda · 1997
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Derivatives of probability measures-concepts and applications to the optimization of stochastic systems
Georg Ch Pflug · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey · 2010
Earlier work this paper cites.
Nonparametric bayesian inverse reinforcement learning for multiple reward functions
Jaedeug Choi and Kee-Eung Kim · 2012
Earlier work this paper cites.
Markov chains and stochastic stability
Sean P Meyn and Richard L Tweedie · 2012
Earlier work this paper cites.
Hierarchical bayesian inverse reinforcement learning
Jaedeug Choi and Kee-Eung Kim · 2014
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Obtaining reward functions of rats using inverse reinforcement learning
Can Eren Sezener, Eiji Uchibe, and Kenji Doya · 2014
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Robust imitation of diverse behaviors
Ziyu Wang, Josh S Merel, Scott E Reed, Nando de Freitas, Gregory Wayne, and Nicolas Heess · 2017
Cited alongside, same era.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
Can ai predict animal movements? filling gaps in animal trajectories using inverse reinforcement learning
Tsubasa Hirakawa, Takayoshi Yamashita, Toru Tamaki, Hironobu Fujiyoshi, Yuta Umezu, Ichiro Takeuchi, Sakiko Matsumoto, and Ken Yoda · 2018
Cited alongside, same era.
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon · 2021
Later among the works it cites.
Finite-sample analysis of off-policy natural actor-critic algorithm
Sajad Khodadadian, Zaiwei Chen, and Siva Theja Maguluri · 2021
Later among the works it cites.
f-irl: Inverse reinforcement learning via state marginal matching
Tianwei Ni, Harshit Sikchi, Yufei Wang, Tejus Gupta, Lisa Lee, and Ben Eysenbach · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Task-relevant adversarial imitation learning
Konrad Zolna, Scott Reed, Alexander Novikov, Sergio Gomez Colmenarejo, David Budden, Serkan Cabi, Misha Denil, Nando de Freitas, and Ziyu Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inverse reinforcement learning of bird flocking behavior
Robert Pinsler, Max Maag, Oleg Arenz, and Gerhard Neumann · 2018
Cited alongside, same era.
Identification of animal behavioral strategies by inverse reinforcement learning
Shoichiro Yamaguchi, Honda Naoki, Muneki Ikeda, Yuki Tsukada, Shunji Nakano, Ikue Mori, and Shin Ishii · 2018
Cited alongside, same era.
A note on the linear convergence of policy gradient methods
Jalaj Bhandari and Daniel Russo · 2020
Cited alongside, same era.
Improving sample complexity bounds for (natural) actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2020
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Cited alongside, same era.
Scalable bayesian inverse reinforcement learning
Alex J Chan and Mihaela van der Schaar · 2021
Cited alongside, same era.
Guanghui Lan · 2022
Later among the works it cites.
Stochastic first-order methods for average-reward markov decision processes
Tianjiao Li, Feiyang Wu, and Guanghui Lan · 2022
Later among the works it cites.
Maximum-likelihood inverse reinforcement learning with finite-time guarantees
Siliang Zeng, Chenliang Li, Alfredo Garcia, and Mingyi Hong · 2022
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2023
Closest in time.
Towards theoretical understanding of inverse reinforcement learning
Alberto Maria Metelli, Filippo Lazzati, and Marcello Restelli · 2023
Closest in time.
Siliang Zeng, Chenliang Li, Alfredo Garcia, and Mingyi Hong · 2023
Closest in time.