Fetching the paper…
Reading the bibliography…
Applying reinforcement learning in physical-world tasks is extremely challenging.
Gait and the energetics of locomotion in horses
Hoyt, D. F. and Taylor, C. R · 1981
Earlier work this paper cites.
A mechanical trigger for the trot–gallop transition in horses
Farley, C. T. and Taylor, C. R · 1991
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1992
Earlier work this paper cites.
Learning agents for uncertain environments
Russell, S · 1998
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, D., and Bagnell, D · 2011
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Machine learning applications for data center optimization
Gao, J. and Jamidar, R · 2014
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M, Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Model-based adversarial imitation learning
Baram, N., Anschel, O., and Mannor, S · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
On-line active reward learning for policy optimisation in spoken dialogue systems
Su, P., Gasic, M., Mrksic, N., Rojas-Barahona, L., Ultes, S., Vandyke, D., Wen, T., and Young, S · 2016
Later among the works it cites.
Duan, Y., Andrychowicz, M., Stadie, B., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Socially compliant mobile robot navigation via inverse reinforcement learning
Kretzschmar, H., Spies, M., Sprunk, C., and Burgard, W · 2016
Cited alongside, same era.
Imitating driver behavior with generative adversarial networks
Kuefler, A., Morton, J., Wheeler, T., and Kochenderfer, M
Cited in the paper.
Stadie, B. C., Abbeel, P., and Sutskever, I · 2017
Later among the works it cites.
Stablizing reinforcement learning in dynamic environment with application to online recommendation
Chen, S.-Y., Yu, Y., Da, Q., Tan, J., Huang, H.-K., and Tang, H.-H · 2018
Closest in time.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Hu, Y., Da, Q., Zeng, A., Yu, Y., and Xu, Y · 2018
Closest in time.