Fetching the paper…
Reading the bibliography…
Policies for partially observed Markov decision processes can be efficiently learned by imitating policies for the corresponding fully observed Markov decision processes.
An analysis of stochastic shortest path problems
Bertsekas, D. P. and Tsitsiklis, J. N · 1991
Earlier work this paper cites.
Reinforcement Learning
Sutton, R · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
Littman, M. L., Cassandra, A. R., and Kaelbling, L. P · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
A survey of POMDP solution techniques
Murphy, K. P · 2000
Earlier work this paper cites.
Asymmetric multiagent reinforcement learning
Könönen, V · 2004
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2004
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Maei, H. R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. S · 2009
Earlier work this paper cites.
A new learning paradigm: Learning using privileged information
Vapnik, V. and Vashist, A · 2009
Earlier work this paper cites.
Approximate policy iteration: A survey and some new methods
Bertsekas, D. P · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, J. A · 2011
Earlier work this paper cites.
Partially Observable Markov Decision Processes , pp. 387–414
Spaan, M. T. J · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., Peters, J., et al · 2013
Earlier work this paper cites.
Bayesian nonparametric methods for partially-observable reinforcement learning
Doshi-Velez, F., Pfau, D., Wood, F., and Roy, N · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
TORCS: The Open Racing Car Simulator, 2014
Wymann, B., Espie, C. G., Dimitrakakis, C., Coulom, R., and Sumner, A · 2014
Earlier work this paper cites.
An Open Approach to Autonomous Vehicles
Kato, S., Takeuchi, E., Ishiguro, Y., Ninomiya, Y., Takeda, K., and Hamada, T · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
CARLA: An open urban driving simulator
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V · 2017
Truncated horizon policy search: Combining reinforcement learning & imitation learning
Sun, W., Bagnell, J. A., and Boots, B · 2018
Later among the works it cites.
Reinforcement learning and optimal control
Bertsekas, D. P · 2019
Later among the works it cites.
Conditional teacher-student learning
Meng, Z., Li, J., Zhao, Y., and Gong, Y · 2019
Later among the works it cites.
Attention-privileged reinforcement learning
Salter, S., Rao, D., Wulfmeier, M., Hadsell, R., and Posner, I · 2019
Later among the works it cites.
Simultaneously learning vision and feature-based control policies for real-world ball-in-a-cup
Schwab, D., Springenberg, J. T., Martins, M. F., Neunert, M., Lampe, T., Abdolmaleki, A., Hertweck, T., Hafner, R., Nori, F., and Riedmiller, M. A · 2019
Later among the works it cites.
Co-training for policy learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dart: Noise injection for robust imitation learning
Laskey, M., Lee, J., Fox, R., Dragan, A., and Goldberg, K · 2017
Cited alongside, same era.
Asymmetric actor critic for image-based robot learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Deeply AggreVaTeD: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Cited alongside, same era.
Information dropout: Learning optimal representations through noisy computation
Achille, A. and Soatto, S · 2018
Cited alongside, same era.
Arora, S., Choudhury, S., and Scherer, S · 2018
Cited alongside, same era.
Song, J., Lanka, R., Yue, Y., and Ono, M · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2020
Closest in time.
Learning dexterous in-hand manipulation
Andrychowicz, O. A. M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W · 2020
Closest in time.
Experiment Tracking with Weights and Biases, 2020
Biewald, L · 2020
Closest in time.
Learning by cheating
Chen, D., Zhou, B., Koltun, V., and Krähenbühl, P · 2020
Closest in time.
Implementation matters in deep rl: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Closest in time.
Privileged information dropout in reinforcement learning
Kamienny, P.-A., Arulkumaran, K., Behbahani, F., Boehmer, W., and Whiteson, S · 2020
Closest in time.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Closest in time.
Belief-grounded networks for accelerated robot learning under partial observability
Nguyen, H., Daley, B., Song, X., Amato, C., and Platt, R · 2020
Closest in time.
Bridging the imitation gap by adaptive insubordination
Weihs, L., Jain, U., Salvador, J., Lazebnik, S., Kembhavi, A., and Schwing, A · 2020
Closest in time.
Behavioral cloning from noisy demonstrations
Sasaki, F. and Yamashina, R · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2021
Closest in time.