Fetching the paper…
Reading the bibliography…
Many real-world systems problems require reasoning about the long term consequences of actions taken to configure and manage the system.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Bandit problems: sequential allocation of experiments (monographs on statistics and applied probability)
Berry, D. A., and Fristedt, B · 1985
Earlier work this paper cites.
Q-learning
Watkins, C. J., and Dayan, P · 1992
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, L.-J · 1993
Earlier work this paper cites.
Packet routing in dynamically changing networks: A reinforcement learning approach
Boyan, J. A., and Littman, M. L · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A., and Niranjan, M · 1994
Earlier work this paper cites.
Predictive q-routing: A memory-based reinforcement learning approach to adaptive traffic control
Choi, S. P., and Yeung, D.-Y · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1997
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K · 2000
Earlier work this paper cites.
Algorithm selection using reinforcement learning
Lagoudakis, M. G., and Littman, M. L · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
To collect or not to collect? machine learning for memory management
Andreasson, E., Hoffmann, F., and Lindholm, O · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
A taxonomy and survey of grid resource management systems for distributed computing
Krauter, K., Buyya, R., and Maheswaran, M · 2002
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
Peters, J., Vijayakumar, S., and Schaal, S · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
Online performance management using hybrid reinforcement learning
Tesauro, G., Das, R., and Jong, N. K · 2006
Earlier work this paper cites.
A reinforcement learning approach to automatic error recovery
Zhu, Q., and Yuan, C · 2007
Earlier work this paper cites.
Feature selection and policy optimization for distributed instruction placement using reinforcement learning
Coons, K. E., Robatmili, B., Taylor, M. E., Maher, B. A., Burger, D., and McKinley, K. S · 2008
Earlier work this paper cites.
Self-optimizing memory controllers: A reinforcement learning approach
Ipek, E., Mutlu, O., Martínez, J. F., and Caruana, R · 2008
Earlier work this paper cites.
Learning classifier tables for autonomic systems on chip
Zeppenfeld, J., Bouajila, A., Stechele, W., and Herkersdorf, A · 2008
Earlier work this paper cites.
Vconf: a reinforcement learning approach to virtual machines auto-configuration
Rao, J., Bu, X., Xu, C.-Z., Wang, L., and Yin, G · 2009
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Deisenroth, M. P., Rasmussen, C. E., and Fox, D · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Mitigating the compiler optimization phase-ordering problem using machine learning
Kulkarni, S., and Cavazos, J · 2012
Cited alongside, same era.
Url: A unified reinforcement learning approach for autonomic cloud management
Xu, C.-Z., Rao, J., and Bu, X · 2012
Cited alongside, same era.
Applying reinforcement learning towards automating resource allocation and application scalability in the cloud
Barrett, E., Howley, E., and Duggan, J · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Cited alongside, same era.
A distributed reinforcement learning scheme for network routing
Littman, M., and Boyan, J · 2013
Ray rllib: A composable and scalable reinforcement learning library
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Gonzalez, J., Goldberg, K., and Stoica, I · 2017
Later among the works it cites.
A hierarchical framework of cloud resource allocation and power management using deep reinforcement learning
Liu, N., Li, Z., Xu, J., Xu, Z., Lin, S., Qiu, Q., Tang, J., and Wang, Y · 2017
Later among the works it cites.
Optimal and scalable caching for 5g using reinforcement learning of space-time popularities
Sadeghi, A., Sheikholeslami, F., and Giannakis, G. B · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
A deep reinforcement learning based framework for power-efficient resource allocation in cloud rans
Xu, Z., Wang, Y., Tang, J., Wang, J., and Gursoy, M. C · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Reinforcement learning-based inter-and intra-application thermal optimization for lifetime improvement of multicore systems
Das, A., Shafik, R. A., Merrett, G. V., Al-Hashimi, B. M., Kumar, A., and Veeravalli, B · 2014
Cited alongside, same era.
Self-tuning intel transactional synchronization extensions
Diegues, N., and Romano, P · 2014
Cited alongside, same era.
Energy-efficient virtual machines consolidation in cloud data centers using reinforcement learning
Farahnakian, F., Liljeberg, P., and Plosila, J · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X · 2014
Cited alongside, same era.
Self-learning cloud controllers: Fuzzy q-learning for knowledge evolution
Jamshidi, P., Sharifloo, A. M., Pahl, C., Metzger, A., and Estrada, G · 2015
Cited alongside, same era.
Later among the works it cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., Xiong, C., and Socher, R · 2017
Later among the works it cites.
A survey on compiler autotuning using machine learning
Ashouri, A. H., Killian, W., Cavazos, J., Palermo, G., and Silvano, C · 2018
Later among the works it cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Later among the works it cites.
Horizon: Facebook’s open source applied reinforcement learning platform
Gauci, J., Conti, E., Liang, Y., Virochsiri, K., He, Y., Kaden, Z., Narayanan, V., and Ye, X · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Later among the works it cites.
Learning to optimize join queries with deep reinforcement learning
Krishnan, S., Yang, Z., Goldberg, K., Hellerstein, J., and Stoica, I · 2018
Later among the works it cites.
Machine learning for internet of things data analysis: A survey
Mahdavinejad, M. S., Rezvan, M., Barekatain, M., Adibi, P., Barnaghi, P., and Sheth, A. P · 2018
Later among the works it cites.
Deep reinforcement learning for join order enumeration
Marcus, R., and Papaemmanouil, O · 2018
Later among the works it cites.
Reinforcement-learning-based foresighted task scheduling in cloud computing
Mostafavi, S., Ahmadi, F., and Sarram, M. A · 2018
Later among the works it cites.
Learning state representations for query optimization with deep reinforcement learning
Ortiz, J., Balazinska, M., Gehrke, J., and Keerthi, S. S · 2018
Later among the works it cites.
Iroko: A framework to prototype reinforcement learning for data center traffic control
Ruffy, F., Przystupa, M., and Beschastnikh, I · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S., and Barto, A. G · 2018
Later among the works it cites.
Machine learning in compiler optimization
Wang, Z., and O’Boyle, M · 2018
Later among the works it cites.
Placeto: Learning generalizable device placement algorithms for distributed machine learning
Addanki, R., Venkatakrishnan, S. B., Gupta, S., Mao, H., and Alizadeh, M · 2019
Closest in time.
Autophase: Compiler phase-ordering for hls with deep reinforcement learning
Huang, Q., Haj-Ali, A., Moses, W., Xiang, J., Stoica, I., Asanovic, K., and Wawrzynek, J · 2019
Closest in time.
A deep reinforcement learning perspective on internet congestion control
Jay, N., Rotman, N., Godfrey, B., Schapira, M., and Tamar, A · 2019
Closest in time.
Liang, E., Zhu, H., Jin, X., and Stoica, I · 2019
Closest in time.
Applications of deep reinforcement learning in communications and networking: A survey
Luong, N. C., Hoang, D. T., Gong, S., Niyato, D., Wang, P., Liang, Y.-C., and Kim, D. I · 2019
Closest in time.
Park: An open platform for learning augmented computer systems
Mao, H., Negi, P., Narayan, A., Wang, H., Yang, J., Wang, H., Marcus, R., Addanki, R., Khani, M., He, S., et al · 2019
Closest in time.
Neo: A learned query optimizer
Marcus, R., Negi, P., Mao, H., Zhang, C., Alizadeh, M., Kraska, T., Papaemmanouil, O., and Tatbul, N · 2019
Closest in time.
Regal: Transfer learning for fast optimization of computation graphs
Paliwal, A., Gimeno, F., Nair, V., Li, Y., Lubin, M., Kohli, P., and Vinyals, O · 2019
Closest in time.