Fetching the paper…
Reading the bibliography…
We start with a brief introduction to reinforcement learning (RL), about its successful stories, basics, an example, issues, the ICML 2019 Workshop on RL for Real Life, how to use it, study material and an outlook.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Valuing American options by simulation: a simple least-squares approach
Longstaff, F. A. and Schwartz, E. S. (2001) · 2001
Earlier work this paper cites.
Regression methods for pricing complex American-style options
Tsitsiklis, J. N. and Van Roy, B. (2001) · 2001
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Learning exercise policies for American options
Li, Y., Szepesvári, C., and Schuurmans, D. (2009) · 2009
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach (3rd edition)
Russell, S. and Norvig, P. (2009) · 2009
Earlier work this paper cites.
An approximate dynamic programming algorithm for large-scale fleet management: A case application
Simão, H. P., Day, J., George, A. P., Gifford, T., Nienow, J., and Powell, W. B. (2009) · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Adaptive stochastic control for the smart grid
Anderson, R. N., Boulanger, A., Powell, W. B., and Scott, W. (2011) · 2011
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality (2nd Edition)
Powell, W. B. (2011) · 2011
Earlier work this paper cites.
Machine learning for market microstructure and high frequency trading
Kearns, M. and Nevmyvaka, Y. (2013) · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Dynamic treatment regimes
Chakraborty, B. and Murphy, S. A. (2014) · 2014
Earlier work this paper cites.
Options, Futures and Other Derivatives (9th edition)
Hull, J. C. (2014) · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
High-confidence off-policy evaluation
Thomas, P. S., Theocharous, G., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Making contextual decisions with low technical debt
Agarwal, A., Bird, S., Cozowicz, M., Hoang, L., Langford, J., Lee, S., Li, J., Melamed, D., Oshri, G., Ribas, O., Sen, S., and Slivkins, A. (2016) · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L. (2016) · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Harley, T., Lillicrap, T. P., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T. (2017) · 2017
Cited alongside, same era.
Learning combinatorial optimization algorithms over graphs
Deep reinforcement learning for de novo drug design
Popova, M., Isayev, O., and Tropsha, A. (2018) · 2018
Later among the works it cites.
Individualized sepsis treatment using reinforcement learning
Saria, S. (2018) · 2018
Later among the works it cites.
Planning chemical syntheses with deep neural networks and symbolic AI
Segler, M. H. S., Preuss, M., and Waller, M. P. (2018) · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018) · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction (2nd Edition)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Intellilight: A reinforcement learning approach for intelligent traffic light control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai, H., Khalil, E. B., Zhang, Y., Dilkina, B., and Song, L. (2017) · 2017
Cited alongside, same era.
Deep Reinforcement Learning: An Overview
Li, Y. (2017) · 2017
Cited alongside, same era.
Device placement optimization with reinforcement learning
Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., and Mohammad Norouzi, Samy Bengio, J. D. (2017) · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisý, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M. (2017) · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017) · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V. (2017) · 2017
Cited alongside, same era.
Wei, H., Zheng, G., Yao, H., and Li, Z. (2018) · 2018
Later among the works it cites.
Deep learning based recommender system: A survey and new perspectives
Zhang, S., Yao, L., Sun, A., and Tay, Y. (2018) · 2018
Later among the works it cites.
DRN: A deep reinforcement learning framework for news recommendation
Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N. J., Xie, X., and Li, Z. (2018) · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. (2018) · 2018
Later among the works it cites.
Software engineering for machine learning: A case study
Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., and Zimmermann, T. (2019) · 2019
Closest in time.
Reinforcement Learning and Optimal Control (draft)
Bertsekas, D. P. (2019) · 2019
Closest in time.
Generative adversarial user model for reinforcement learning based recommendation system
Chen, X., Li, S., Li, H., Jiang, S., Qi, Y., and Song, L. (2019) · 2019
Closest in time.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. (2019) · 2019
Closest in time.
Horizon: Facebook’s open source applied reinforcement learning platform
Gauci, J., Conti, E., Liang, Y., Virochsiri, K., He, Y., Kaden, Z., Narayanan, V., Ye, X., and Chen, Z. (2019) · 2019
Closest in time.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., and Celi, L. A. (2019) · 2019
Closest in time.
Coloring big graphs with alphagozero
Huang, J., Patwary, M., and Diamos, G. (2019) · 2019
Closest in time.
Learning agile and dynamic motor skills for legged robots
Hwangbo, J., Lee, J., Dosovitskiy, A., Bellicoso, D., Tsounis, V., Koltun, V., and Hutter, M. (2019) · 2019
Closest in time.
Lessons from real-world reinforcement learning in a customer support bot
Karampatziakis, N., Kochman, S., Huang, J., Mineiro, P., Osborne, K., and Chen, W. (2019) · 2019
Closest in time.
Attention, learn to solve routing problems!
Kool, W., van Hoof, H., and Welling, M. (2019) · 2019
Closest in time.
Efficient ridesharing order dispatching with mean field multi-agent reinforcement learning
Li, M., (Tony)Qin, Z., Jiao, Y., Yang, Y., Gong, Z., Wang, J., Wang, C., Wu, G., and Ye, J. (2019) · 2019
Closest in time.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Shi, J.-C., Yu, Y., Da, Q., Chen, S.-Y., and Zeng, A.-X. (2019) · 2019
Closest in time.
A deep value-network based approach for multi-driver order dispatching
Tang, X., Qin, Z., Zhang, F., Wang, Z., Xu, Z., Ma, Y., Zhu, H., and Ye, J. (2019) · 2019
Closest in time.
Reinforcement learning for online information seeking
Zhao, X., Xia, L., Tang, J., and Yin, D. (2019) · 2019
Closest in time.