Fetching the paper…
Reading the bibliography…
An increasingly important building block of large scale machine learning systems is based on returning slates; an ordered lists of items given a query.
Accelerating large-scale inference with anisotropic vector quantization
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar · 1908
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
The analysis of permutations
R. L. Plackett · 1975
Earlier work this paper cites.
Evolutionsstrategien
I. Rechenberg · 1978
Earlier work this paper cites.
The singular value decomposition: Its computation and some applications
V. Klema and A. Laub · 1980
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Aristides Gionis, Piotr Indyk, and Rajeev Motwani · 1999
Earlier work this paper cites.
Listwise approach to learning to rank: Theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li · 2008
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu · 2009
Earlier work this paper cites.
Bpr: Bayesian personalized ranking from implicit feedback
Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme · 2009
Earlier work this paper cites.
Variational optimization, 2012
Joe Staines and David Barber · 2012
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Monte Carlo theory, methods and examples
Art B. Owen · 2013
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Anshumali Shrivastava and Ping Li · 2014
Earlier work this paper cites.
The movielens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2015
Cited alongside, same era.
Counterfactual risk minimization: Learning from logged bandit feedback
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Randomized Quasi-Monte Carlo: An Introduction for Practitioners
Pierre L’Ecuyer · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Meta-prod2vec: Product embeddings using side-information for recommendation
Flavian Vasile, Elena Smirnova, and Alexis Conneau · 2016
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford, Damien Jose, and Imed Zitouni · 2017
Fine-grained spoiler detection from large-scale review corpora
Mengting Wan, Rishabh Misra, Ndapa Nakashole, and Julian J. McAuley · 2019
Later among the works it cites.
Low-variance black-box gradient estimates for the plackett-luce distribution
Artyom Gadetsky, Kirill Struminsky, Christopher Robinson, Novi Quadrianto, and Dmitry Vetrov · 2020
Later among the works it cites.
Off-policy learning in two-stage recommender systems
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, and Ed H. Chi · 2020
Later among the works it cites.
Softsort: A continuous relaxation for the argsort operator
Sebastian Prillo and Julian Martin Eisenschlos · 2020
Later among the works it cites.
BLOB: A Probabilistic Model for Recommendation That Combines Organic and Bandit Signals
Otmane Sakhi, Stephen Bonner, David Rohde, and Flavian Vasile · 2020
Later among the works it cites.
On the convergence of sgd with biased gradients, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud · 2018
Cited alongside, same era.
Variational autoencoders for collaborative filtering
Dawen Liang, Rahul G. Krishnan, Matthew D. Hoffman, and Tony Jebara · 2018
Cited alongside, same era.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Yu A. Malkov and D. A. Yashunin · 2018
Cited alongside, same era.
Item recommendation on monotonic behavior chains
Mengting Wan and Julian J. McAuley · 2018
Cited alongside, same era.
The lambdaloss framework for ranking metric optimization
Xuanhui Wang, Cheng Li, Nadav Golbandi, Mike Bendersky, and Marc Najork · 2018
Cited alongside, same era.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H. Chi · 2019
Cited alongside, same era.
Ahmad Ajalloeian and Sebastian U. Stich · 2021
Later among the works it cites.
Black-Box Optimization: Methods and Applications , pp. 35–65
Ishan Bajaj, Akhil Arora, and M. M. Faruque Hasan · 2021
Later among the works it cites.
A review of the gumbel-max trick and its extensions for discrete stochasticity in machine learning
Iris AM Huijben, Wouter Kool, Max B Paulus, and Ruud JG van Sloun · 2021
Later among the works it cites.
Scalable representation learning and retrieval for display advertising
Olivier Koch, Amine Benhalloum, Guillaume Genthial, Denis Kuzin, and Dmitry Parfenchik · 2021
Later among the works it cites.
Learning-to-rank with partitioned preference: Fast estimation for the plackett-luce model
Jiaqi Ma, Xinyang Yi, Weijing Tang, Zhe Zhao, Lichan Hong, Ed Chi, and Qiaozhu Mei · 2021
Later among the works it cites.
Recommendation on Live-Streaming Platforms: Dynamic Availability and Repeat Consumption , pp. 390–399
Jérémie Rappaz, Julian McAuley, and Karl Aberer · 2021
Later among the works it cites.
Low-variance estimation in the plackett-luce model via quasi-monte carlo sampling, 2022
Alexander Buchholz, Jan Malte Lichtenberg, Giuseppe Di Benedetto, Yannik Stein, Vito Bellini, and Matteo Ruffini · 2022
Later among the works it cites.
Off-policy actor-critic for recommender systems
Minmin Chen, Can Xu, Vince Gatto, Devanshu Jain, Aviral Kumar, and Ed Chi · 2022
Later among the works it cites.
Reward shaping for user satisfaction in a reinforce recommender, 2022
Konstantina Christakopoulou, Can Xu, Sai Zhang, Sriraj Badam, Trevor Potter, Daniel Li, Hao Wan, Xinyang Yi, Ya Le, Chris Berg, Eric Bencomo Dixon, Ed H. Chi, and Minmin Chen · 2022
Later among the works it cites.
Learning-to-rank at the speed of sampling: Plackett-luce gradient estimation with minimal computational complexity
Harrie Oosterhuis · 2022
Later among the works it cites.
Surrogate for long-term user experience in recommender systems
Yuyan Wang, Mohit Sharma, Can Xu, Sriraj Badam, Qian Sun, Lee Richardson, Lisa Chung, Ed H. Chi, and Minmin Chen · 2022
Later among the works it cites.