Fetching the paper…
Reading the bibliography…
Learning from Observations (LfO) is a practical reinforcement learning scenario from which many applications can benefit through the reuse of incomplete resources.
Asymptotic evaluation of certain markov process expectations for large time, i
Monroe D Donsker and SR Srinivasa Varadhan · 1975
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
High speed obstacle avoidance using monocular vision and reinforcement learning
Jeff Michels, Ashutosh Saxena, and Andrew Y Ng · 2005
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
f-gan: Training generative neural samplers using variational divergence minimization
Sebastian Nowozin, Botond Cseke, and Ryota Tomioka · 2016
Earlier work this paper cites.
Deep direct reinforcement learning for financial signal representation and trading
Yue Deng, Feng Bao, Youyong Kong, Zhiquan Ren, and Qionghai Dai · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
Ahmad EL Sallab, Mohammed Abdou, Etienne Perot, and Senthil Yogamani · 2017
Cited alongside, same era.
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
Third-person imitation learning
Bradly C Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Imitation learning from observations by minimizing inverse dynamics disagreement
Chao Yang, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Huaping Liu, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Recent advances in imitation learning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2019
Later among the works it cites.
Sample efficient imitation learning for continuous control
Fumihiro Sasaki, Tetsuya Yohira, and Atsuo Kawaguchi · 2019
Later among the works it cites.
Wasserstein adversarial imitation learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 2019
Later among the works it cites.
State alignment-based imitation learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, and Hao Su · 2019
Later among the works it cites.
Adversarial imitation learning from state-only demonstrations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yusuf Aytar, Tobias Pfaff, David Budden, Thomas Paine, Ziyu Wang, and Nando de Freitas · 2018
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Van Hoof, and David Meger · 2018
Cited alongside, same era.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Cited alongside, same era.
Faraz Torabi, Garrett Warnell, and Peter Stone · 2019
Later among the works it cites.
Provably efficient imitation learning from observation alone
Wen Sun, Anirudh Vemula, Byron Boots, and J Andrew Bagnell · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Later among the works it cites.
Understanding the relation between maximum-entropy inverse reinforcement learning and behaviour cloning
Seyed Kamyar Seyed Ghasemipour, Shane Gu, and Richard Zemel · 2019
Later among the works it cites.
Imitating latent policies from observation
Ashley D Edwards, Himanshu Sahni, Yannick Schroecker, and Charles L Isbell · 2019
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
State-only imitation with transition dynamics mismatch
Tanmay Gangwani and Jian Peng · 2020
Later among the works it cites.