Fetching the paper…
Reading the bibliography…
We consider the dynamic assortment optimization problem under the multinomial logit model (MNL) with unknown utility parameters.
Huber PJ (1964) Robust estimation of a location parameter. The Annals of Mathematical Statistics 35(1):73–101
1964
Earlier work this paper cites.
McFadden D (1974) Conditional logit analysis of qualitative choice behavior. Frontiers in Econometrics (Academic Press)
1974
Earlier work this paper cites.
Freedman DA (1975) On tail probabilities for martingales. The Annals of Probability 3(1):100–118
1975
Earlier work this paper cites.
van Ryzin G, Mahajan S (1999) On the relationships between inventory costs and variety benefits in retail assortments. Management Science 45:1496–1509
1999
Earlier work this paper cites.
Mahajan S, van Ryzin G (2001) Stocking retail assortments under dynamic consumer substitution. Operations Research 49:334–351
2001
Earlier work this paper cites.
Auer P (2002) Using confidence bounds for exploitation-exploration trade-offs. Journal of Machine Learning Research 3(Nov):397–422
2002
Earlier work this paper cites.
Auer P, Cesa-Bianchi N, Freund Y, Schapire RE (2002) The nonstochastic multiarmed bandit problem. SIAM Journal on Computing 32(1):48–77
2002
Earlier work this paper cites.
Cooper WL, de Mello TH, Kleywegt AJ (2006) Models of the spiral-down effect in revenue management. Operations Research 54(5):968–987
2006
Earlier work this paper cites.
Even-Dar E, Mannor S, Mansour Y (2006) Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems. Journal of Machine Learning Research 7(Jun):1079–1105
2006
Earlier work this paper cites.
Caro F, Gallien J (2007) Dynamic Assortment with Demand Learning for Seasonal Consumer Goods. Management Science 53(2):276–292
2007
Earlier work this paper cites.
Auer P, Ortner R (2010) Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem. Periodica Mathematica Hungarica 61(1-2):55–65
2010
Cited alongside, same era.
Rusmevichientong P, Shen ZJ, Shmoys D (2010) Dynamic assortment optimization with a multinomial logit choice model and capacity constraint. Operations Research 58(6):1666–1680
2010
Cited alongside, same era.
Huber PJ, Ronchetti EM (2011) Robust Statistics . Series in Probability and Statistics (Wiley)
2011
Cited alongside, same era.
Bubeck S, Cesa-Bianchi N (2012) Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends in Machine Learning 5(1):1–122
2012
Cited alongside, same era.
Saure D, Zeevi A (2013) Optimal dynamic assortment planning with demand learning. Manufacturing & Service Operations Management 15(3):387–404
2013
Diakonikolas I, Kamath G, Kane DM, Li J, Moitra A, Stewart A (2017) Being robust (in high dimensions) can be practical. Proceedings of the Interational Conference on Machine Learning
2017
Later among the works it cites.
Chen X, Wang Y (2018) A note on tight lower bound for mnl-bandit assortment selection models. Operations Research Letters 46(5):534–537
2018
Later among the works it cites.
2018
Later among the works it cites.
Diakonikolas I, Kamath G, Kane D, Li J, Moitra A, Stewart A (2018) Robustly learning a gaussian: Getting optimal error, efficiently. Proceedings of the ACM-SIAM Symposium on Discrete Algorithms
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Besbes O, Zeevi A (2015) On the (surprising) sufficiency of linear models for dynamic pricing with demand learning. Management Science 61(4):723–739
2015
Cited alongside, same era.
Chen M, Gao C, Ren Z (2016) A general decision theory for huber’s ϵ \epsilon -contamination model. Electronic Journal of Statistics 10(2):3752–3774
2016
Cited alongside, same era.
Agrawal S, Avandhanula V, Goyal V, Zeevi A (2017) Thompson sampling for MNL-bandit. Proccedings of the Conference on Learning Theory (COLT)
2017
Cited alongside, same era.
Cheung WC, Simchi-Levi D (2017) Thompson sampling for online personalized assortment optimization problems with multinomial logit choice models. Available at SSRN: https://papers.ssrn.com/?abstract_id=3075658
2017
Cited alongside, same era.
Esfandiari H, Korula N, Mirrokni V (32018) Allocation with traffic spikes: Mixing adversarial and stochastic models. ACM Transactions on Economics and Computation 6(3–4):1–23
Cited in the paper.
2018
Later among the works it cites.
Wang Y, Chen X, Zhou Y (2018) Near-optimal policies for dynamic multinomial logit assortment selection models. Proceedings of Advances in Neural Information Processing Systems (NeurIPS)
2018
Later among the works it cites.
Agrawal S, Avadhanula V, Goyal V, Zeevi A (2019) MNL-bandit: A dynamic learning approach to assortment selection. Operations Research 67(5):1453–1485
2019
Closest in time.
Gupta A, Koren T, Talwar K (2019) Better algorithms for stochastic bandits with adversarial corruptions. Proceedings of the Conference on Learning Theory
2019
Closest in time.
Oh MH, Iyengar G (2019) Multinomial logit contextual bandits. Reinforcement Learning for Real Life (RL4RealLife) Workshop in the International Conference on Machine Learning (ICML)
2019
Closest in time.