Fetching the paper…
Reading the bibliography…
Many resource management problems require sequential decision-making under uncertainty, where the only uncertainty affecting the decision outcomes are exogenous variables outside the control of the decision-maker.
Forecasting and control of passenger bookings
Littlewood, K · 1972
Earlier work this paper cites.
Optimal search for the best alternative
Weitzman, M. L · 1979
Earlier work this paper cites.
Scenarios and policy aggregation in optimization under uncertainty
Rockafellar, R. T. and Wets, R. J.-B · 1991
Earlier work this paper cites.
Introduction to linear optimization , volume 6
Bertsimas, D. and Tsitsiklis, J. N · 1997
Earlier work this paper cites.
A framework for simulation-based network control via hindsight optimization
Chong, E. K., Givan, R. L., and Chang, H. S · 2000
Earlier work this paper cites.
Pattern classification
Hart, P. E., Stork, D. G., and Duda, R. O · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Pegasus: a policy search method for large mdps and pomdps
Ng, A. Y. and Jordan, M · 2000
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Abbeel, P. and Ng, A. Y · 2005
Earlier work this paper cites.
Limits of dense graph sequences
Lovász, L. and Szegedy, B · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, J. and Zhang, T · 2007
Earlier work this paper cites.
Performance analysis of online anticipatory algorithms for large multistage stochastic integer programs
Mercier, L. and Van Hentenryck, P · 2007
Earlier work this paper cites.
Convergent sequences of dense graphs I: Subgraph frequencies, metric properties and testing
Borgs, C., Chayes, J. T., Lovász, L., Sós, V. T., and Vesztergombi, K · 2008
Earlier work this paper cites.
Secretary problems and incentives via linear programming
Buchbinder, N., Jain, K., and Singh, M · 2009
Earlier work this paper cites.
Gopalan, P., Klivans, A., and Meka, R · 2010
Earlier work this paper cites.
Heuristics for vector bin packing
Panigrahy, R., Talwar, K., Uyeda, L., and Wieder, U · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S. and Cesa-Bianchi, N · 2012
Earlier work this paper cites.
Gupta, V. and Radovanovic, A · 2012
Earlier work this paper cites.
The stochastic generalized bin packing problem
Perboli, G., Tadei, R., and Baldi, M. M · 2012
Earlier work this paper cites.
An infinite server system with general packing constraints
Stolyar, A. L · 2013
Earlier work this paper cites.
Information relaxations, duality, and convex stochastic dynamic programs
Brown, D. B. and Smith, J. E · 2014
Earlier work this paper cites.
Integer programming , volume 271
Conforti, M., Cornuéjols, G., and Zambelli, G · 2014
Earlier work this paper cites.
Asymptotic optimality of bestfit for stochastic bin packing
Ghaderi, J., Zhong, Y., and Srikant, R · 2014
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
Dynamic bin packing for on-demand cloud resource allocation
Li, Y., Tang, X., and Cai, W · 2015
Earlier work this paper cites.
Pointer networks
Vinyals, O., Fortunato, M., and Jaitly, N · 2015
Earlier work this paper cites.
Optimality conditions for inventory control
Feinberg, E. A · 2016
Earlier work this paper cites.
Causal bandits: learning good interventions via causal inference
Lattimore, F., Lattimore, T., and Reid, M. D · 2016
Earlier work this paper cites.
Resource management with deep reinforcement learning
Mao, H., Alizadeh, M., Menache, I., and Kandula, S · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Allen, C., Asadi, K., Roderick, M., Mohamed, A.-r., Konidaris, G., and Littman, M · 2017
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D · 2017
Cited alongside, same era.
Neural combinatorial optimization with reinforcement learning
Bello, I., Pham, H., Le, Q. V., Norouzi, M., and Bengio, S · 2017
Cited alongside, same era.
Information relaxation bounds for infinite horizon markov decision processes
Brown, D. B. and Haugh, M. B · 2017
Cited alongside, same era.
Resource central: Understanding and predicting workloads for improved resource management in large cloud platforms
Cortez, E., Bonde, A., Muzio, A., Russinovich, M., Fontoura, M., and Bianchini, R · 2017
Cited alongside, same era.
Solving a new 3d bin packing problem with deep reinforcement learning method
Hu, H., Zhang, X., Yan, X., Wang, L., and Xu, Y · 2017
Cited alongside, same era.
PowerNet: Multi-agent deep reinforcement learning for scalable powergrid control
Chen, D., Chen, K., Li, Z., Chu, T., Yao, R., Qiu, F., and Lin, K · 2021
Later among the works it cites.
Heuristic-guided reinforcement learning
Cheng, C.-A., Kolobov, A., and Swaminathan, A · 2021
Later among the works it cites.
Queueing network controls via deep reinforcement learning
Dai, J. G. and Gluzman, M · 2021
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Domingues, O. D., Ménard, P., Kaufmann, E., and Valko, M · 2021
Later among the works it cites.
Universal trading for order execution with oracle policy distillation
Fang, Y., Ren, K., Liu, W., Zhou, D., Zhang, W., Bian, J., Yu, Y., and Liu, T.-Y · 2021
Later among the works it cites.
Scalable deep reinforcement learning for ride-hailing
Feng, J., Gluzman, M., and Dai, J. G · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Cited alongside, same era.
Discovering and Removing Exogenous State Variables and Rewards for Reinforcement Learning
Dietterich, T., Trimponias, G., and Chen, Z · 2018
Cited alongside, same era.
Improving online algorithms via ml predictions
Kumar, R., Purohit, M., and Svitkina, Z · 2018
Cited alongside, same era.
A tutorial on thompson sampling
Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., and Wen, Z · 2018
Cited alongside, same era.
Learning to search via retrospective imitation
Song, J., Lanka, R., Zhao, A., Bhatnagar, A., Yue, Y., and Ono, M · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Later among the works it cites.
The statistical complexity of interactive decision making
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A · 2021
Later among the works it cites.
A Survey of Recent Progress in the Asymptotic Analysis of Inventory Systems
Goldberg, D. A., Reiman, M. I., and Wang, Q · 2021
Later among the works it cites.
Math programming based reinforcement learning for multi-echelon inventory management
Harsha, P., Jagmohan, A., Kalagnanam, J., Quanz, B., and Singhvi, D · 2021
Later among the works it cites.
Exploiting action impact regularity and partially known models for offline reinforcement learning
Liu, V., Wright, J., and White, M · 2021
Later among the works it cites.
Competitive caching with machine learned advice
Lykouris, T. and Vassilvitskii, S · 2021
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Mesnard, T., Weber, T., Viola, F., Thakoor, S., Saade, A., Harutyunyan, A., Dabney, W., Stepleton, T. S., Heess, N., Guez, A., Moulines, E., Hutter, M., Buesing, L., and Munos, R · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Later among the works it cites.
VMAgent: Scheduling simulator for reinforcement learning
Sheng, J., Cai, S., Cui, H., Li, W., Hua, Y., Jin, B., Zhou, W., Hu, Y., Zhu, L., Peng, Q., Zha, H., and Wang, X · 2021
Later among the works it cites.
The bayesian prophet: A low-regret framework for online decision making
Vera, A. and Banerjee, S · 2021
Later among the works it cites.
Online allocation and pricing: Constant regret via bellman inequalities
Vera, A., Banerjee, S., and Gurvich, I · 2021
Later among the works it cites.
Robust asymmetric learning in POMDPs
Warrington, A., Lavington, J. W., Scibior, A., Schmidt, M., and Wood, F · 2021
Later among the works it cites.
Explaining fast improvement in online imitation learning
Yan, X., Boots, B., and Cheng, C.-A · 2021
Later among the works it cites.
Learning in structured mdps with convex cost functions: Improved regret bounds for inventory management
Agrawal, S. and Jia, R · 2022
Closest in time.
Orsuite: Benchmarking suite for sequential operations models
Archer, C., Banerjee, S., Cortez, M., Rucker, C., Sinclair, S. R., Solberg, M., Xie, Q., and Lee Yu, C · 2022
Closest in time.
Outcome-driven dynamic refugee assignment with allocation balancing
Bansak, K. and Paulson, E · 2022
Closest in time.
Information relaxations and duality in stochastic dynamic programs:: A review and tutorial
Brown, D. B. and Smith, J. E · 2022
Closest in time.
Adversarially trained actor critic for offline reinforcement learning
Cheng, C.-A., Xie, T., Jiang, N., and Agarwal, A · 2022
Closest in time.
Sample-efficient reinforcement learning in the presence of exogenous information
Efroni, Y., Foster, D. J., Misra, D., Krishnamurthy, A., and Langford, J · 2022
Closest in time.
Smart “predict, then optimize”
Elmachtoub, A. N. and Grigas, P · 2022
Closest in time.
Stateful offline contextual policy evaluation and learning
Kallus, N. and Zhou, A · 2022
Closest in time.
Efficient reinforcement learning with prior causal knowledge
Lu, Y., Meisami, A., and Tewari, A · 2022
Closest in time.
Madeka, D., Torkkola, K., Eisenach, C., Foster, D., and Luo, A · 2022
Closest in time.
Reinforcement Learning and Stochastic Optimization: A unified framework for sequential decisions
Powell, W · 2022
Closest in time.
Learning to schedule multi-NUMA virtual machines via reinforcement learning
Sheng, J., Hu, Y., Zhou, W., Zhu, L., Jin, B., Wang, J., and Wang, X · 2022
Closest in time.
Policy gradients incorporating the future
Venuto, D., Lau, E., Precup, D., and Nachum, O · 2022
Closest in time.
Distributionally robust inventory control when demand is a martingale
Xin, L. and Goldberg, D. A · 2022
Closest in time.