Fetching the paper…
Reading the bibliography…
Approaches to recommendation are typically evaluated in one of two ways: (1) via a (simulated) online experiment, often seen as the gold standard, or (2) via some offline evaluation procedure, where the goal is to approximate the outcome of an online experiment.
On the Difficulty of Evaluating Baselines: A Study on Recommender Systems
S. Rendle, L. Zhang, and Y. Koren. 2019 · 1905
Earlier work this paper cites.
On the Value of Bandit Feedback for Offline Recommender System Evaluation
O. Jeunen, D. Rohde, and F. Vasile. 2019 · 1907
Earlier work this paper cites.
Probable Error of a Correlation Coefficient
Student. 1908 · 1908
Earlier work this paper cites.
RecSim: A Configurable Simulation Platform for Recommender Systems
E. Ie, C. Hsu, M. Mladenov, V. Jain, S. Narvekar, J. Wang, R. Wu, and C. Boutilier. 2019a · 1909
Earlier work this paper cites.
Improvements That Don’t Add up: Ad-Hoc Retrieval Results since 1998. In Proc of. the 18th ACM Conference on Information and Knowledge Management (CIKM ’09) . ACM, 601–610
T. G. Armstrong, A. Moffat, W. Webber, and J. Zobel. 2009 · 1998
Earlier work this paper cites.
Cumulated Gain-Based Evaluation of IR Techniques
K. Järvelin and J. Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Evaluating Collaborative Filtering Recommender Systems
J. L. Herlocker, J. A. Konstan, L. G. Terveen, and J. T. Riedl. 2004 · 2004
Earlier work this paper cites.
The Relationship between IR Effectiveness Measures and User Satisfaction. In Proc of. the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’07) . ACM, 773–774
A. Al-Maskari, M. Sanderson, and P. Clough. 2007 · 2007
Earlier work this paper cites.
The Netflix prize. In Proc. of the KDD cup and workshop , Vol. 2007. 35
J. Bennett, S. Lanning, et al · 2007
Earlier work this paper cites.
Novelty and Diversity in Information Retrieval Evaluation. In Proc. of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’08) . ACM, 659–666
C.L.A. Clarke, M. Kolla, G. V. Cormack, O. Vechtomova, A. Ashkan, S. Büttcher, and I. MacKinnon. 2008 · 2008
Earlier work this paper cites.
An Experimental Comparison of Click Position-Bias Models. In Proc of. the 2008 International Conference on Web Search and Data Mining (WSDM ’08) . ACM, 87–94
N. Craswell, O. Zoeter, M. Taylor, and B. Ramsey. 2008 · 2008
Earlier work this paper cites.
Truncated Importance Sampling
E. L. Ionides. 2008 · 2008
Earlier work this paper cites.
Rank-Biased Precision for Measurement of Retrieval Effectiveness
A. Moffat and J. Zobel. 2008 · 2008
Earlier work this paper cites.
Introduction to information retrieval . Vol. 39
H. Schütze, C. D Manning, and P. Raghavan. 2008 · 2008
Earlier work this paper cites.
Expected Reciprocal Rank for Graded Relevance. In Proc of. the 18th ACM Conference on Information and Knowledge Management (CIKM ’09) . ACM, 621–630
O. Chapelle, D. Metzler, Y. Zhang, and P. Grinspan. 2009 · 2009
Earlier work this paper cites.
Doubly Robust Policy Evaluation and Learning. In Proc. of the 28th International Conference on International Conference on Machine Learning (ICML’11) . 1097–1104
M. Dudík, J. Langford, and L. Li. 2011 · 2011
Earlier work this paper cites.
A Comparative Analysis of Offline and Online Evaluations and Discussion of Research Paper Recommender System Evaluation. In Proc. of the International Workshop on Reproducibility and Replication in Recommender Systems Evaluation (RepSys ’13) . 7–14
J. Beel, M. Genzmehr, S. Langer, A. Nürnberger, and B. Gipp. 2013 · 2013
Earlier work this paper cites.
Two-Stage Learning to Rank for Information Retrieval. In Advances in Information Retrieval . Springer Berlin Heidelberg, 423–434
V. Dang, M. Bendersky, and W. B. Croft. 2013 · 2013
Earlier work this paper cites.
Monte Carlo theory, methods and examples
A. B. Owen. 2013 · 2013
Earlier work this paper cites.
Evaluation of Recommendations: Rating-prediction and Ranking. In Proc. of the 7th ACM Conference on Recommender Systems (RecSys ’13) . ACM, 213–220
H. Steck. 2013 · 2013
Earlier work this paper cites.
Offline and Online Evaluation of News Recommender Systems at Swissinfo.Ch. In Proc. of the 8th ACM Conference on Recommender Systems (RecSys ’14) . 169–176
F. Garcin, B. Faltings, O. Donatsch, A. Alazzawi, C. Bruttin, and A. Huber. 2014 · 2014
Earlier work this paper cites.
Unifying Nearest Neighbors Collaborative Filtering. In Proc. of the 8th ACM Conference on Recommender Systems (RecSys ’14) . ACM, 177–184
K. Verstrepen and B. Goethals. 2014 · 2014
Earlier work this paper cites.
Click Models for Web Search
A. Chuklin, I. Markov, and M. de Rijke. 2015 · 2015
Earlier work this paper cites.
The MovieLens Datasets: History and Context
F. Maxwell Harper and Joseph A. Konstan. 2015 · 2015
Earlier work this paper cites.
The Self-Normalized Estimator for Counterfactual Learning. In Advances in Neural Information Processing Systems . 3231–3239
A. Swaminathan and T. Joachims. 2015 · 2015
Earlier work this paper cites.
A Neural Click Model for Web Search. In Proc. of the 25th International Conference on World Wide Web (WWW ’16) . 531–541
A. Borisov, I. Markov, M. de Rijke, and P. Serdyukov. 2016 · 2016
Earlier work this paper cites.
Deep Neural Networks for YouTube Recommendations. In Proc. of the 10th ACM Conference on Recommender Systems (RecSys ’16) . ACM, 191–198
P. Covington, J. Adams, and E. Sargin. 2016 · 2016
Earlier work this paper cites.
Contrasting Offline and Online Results when Evaluating Recommendation Algorithms. In Proc. of the 10th ACM Conference on Recommender Systems (RecSys ’16) . ACM, 31–34
M. Rossetti, F. Stella, and M. Zanker. 2016 · 2016
Earlier work this paper cites.
Effective Evaluation Using Logged Bandit Feedback from Multiple Loggers. In Proc. of the 23rd ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’17) . ACM, 687–696
A. Agarwal, S. Basu, T. Schnabel, and T. Joachims. 2017 · 2017
Earlier work this paper cites.
Learning Sensitive Combinations of A/B Test Metrics. In Proc. of the Tenth ACM International Conference on Web Search and Data Mining (WSDM ’17) . ACM, 651–659
E. Kharitonov, A. Drutsa, and P. Serdyukov. 2017 · 2017
Earlier work this paper cites.
Off-policy evaluation for slate recommendation. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc
A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudik, J. Langford, D. Jose, and I. Zitouni. 2017 · 2017
Earlier work this paper cites.
Offline A/B Testing for Recommender Systems. In Proc. of the Eleventh ACM International Conference on Web Search and Data Mining (WSDM ’18) . ACM, 198–206
A. Gilotte, C. Calauzènes, T. Nedelec, A. Abraham, and S. Dollé. 2018 · 2018
Earlier work this paper cites.
Fair Offline Evaluation Methodologies for Implicit-feedback Recommender Systems with MNAR Data. In Proc. of the REVEAL 18 Workshop on Offline Evaluation for Recommender Systems (RecSys ’18)
O. Jeunen, K. Verstrepen, and B. Goethals. 2018 · 2018
Earlier work this paper cites.
Towards a Fair Marketplace: Counterfactual Evaluation of the Trade-off between Relevance, Fairness & Satisfaction in Recommendation Systems. In Proc. of the 27th ACM International Conference on Information and Knowledge Management (CIKM ’18) . ACM, 2243–2251
R. Mehrotra, J. McInerney, H. Bouchard, M. Lalmas, and F. Diaz. 2018 · 2018
Cited alongside, same era.
D. Rohde, S. Bonner, T. Dunlop, F. Vasile, and A. Karatzoglou. 2018 · 2018
Cited alongside, same era.
Unbiased Offline Recommender Evaluation for Missing-not-at-random Implicit Feedback. In Proc. of the 12th ACM Conference on Recommender Systems (RecSys ’18) . ACM, 279–287
L. Yang, Y. Cui, Yuan X., C. Wang, S. Belongie, and D. Estrin. 2018 · 2018
Cited alongside, same era.
Addressing Trust Bias for Unbiased Learning-to-Rank. In Proc. of the 2019 World Wide Web Conference (WWW ’19) . ACM, 4–14
A. Agarwal, X. Wang, C. Li, M. Bendersky, and M. Najork. 2019 · 2019
Cited alongside, same era.
On Evaluating Session-Based Recommendation with Implicit Feedback. In Workshop on Perspectives on Offline Evaluation for Recommender Systems at RecSys ’21 (PERSPECTIVES ’21)
F. Diaz. 2021 · 2021
Later among the works it cites.
Towards Meaningful Statements in IR Evaluation: Mapping Evaluation Measures to Interval Scales
M. Ferrante, N. Ferro, and N. Fuhr. 2021 · 2021
Later among the works it cites.
A Troubling Analysis of Reproducibility and Progress in Recommender Systems Research
M. Ferrari Dacrema, S. Boglio, P. Cremonesi, and D. Jannach. 2021 · 2021
Later among the works it cites.
The Simpson’s Paradox in the Offline Evaluation of Recommendation Systems
A. H. Jadidinejad, C. Macdonald, and I. Ounis. 2021 · 2021
Later among the works it cites.
Recommendations as Treatments
T. Joachims, B. London, Y. Su, A. Swaminathan, and L. Wang. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Top-K Off-Policy Correction for a REINFORCE Recommender System. In Proc. of the 12th ACM International Conference on Web Search and Data Mining (WSDM ’19) . ACM, 456–464
M. Chen, A. Beutel, P. Covington, S. Jain, F. Belletti, and E. H. Chi. 2019 · 2019
Cited alongside, same era.
Generalized Multiple Importance Sampling
V. Elvira, L. Martino, D. Luengo, and M.F. Bugallo. 2019 · 2019
Cited alongside, same era.
Intervention Harvesting for Context-Dependent Examination-Bias Estimation. In Proc of. the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’19) . ACM, 825–834
Z. Fang, A. Agarwal, and T. Joachims. 2019 · 2019
Cited alongside, same era.
Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches. In Proc. of the 13th ACM Conference on Recommender Systems (RecSys ’19) . ACM, 101–109
M. Ferrari Dacrema, P. Cremonesi, and D. Jannach. 2019 · 2019
Cited alongside, same era.
Offline Evaluation to Make Decisions About Playlist Recommendation Algorithms. In Proc of. the 12th ACM International Conference on Web Search and Data Mining (WSDM ’19) . ACM, 420–428
A. Gruson, P. Chandar, C. Charbuillet, J. McInerney, S. Hansen, D. Tardieu, and B. Carterette. 2019 · 2019
Cited alongside, same era.
SlateQ: A tractable decomposition for reinforcement learning with recommendation sets
E. Ie, V. Jain, J. Wang, S. Narvekar, R. Agarwal, R. Wu, H. Cheng, T. Chandra, and C. Boutilier. 2019b · 2019
Cited alongside, same era.
Revisiting Offline Evaluation for Implicit-feedback Recommender Systems. In Proc. of the 13th ACM Conference on Recommender Systems (RecSys ’19) . ACM, 596–600
O. Jeunen. 2019 · 2019
Cited alongside, same era.
Embarrassingly Shallow Autoencoders for Sparse Data. In The World Wide Web Conference (WWW ’19) . ACM, 3251–3257
H. Steck. 2019 · 2019
Cited alongside, same era.
Optimal Off-Policy Evaluation from Multiple Logging Policies. In Proc. of the 38th International Conference on Machine Learning (ICML ’21, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 5247–5256
N. Kallus, Y. Saito, and M. Uehara. 2021 · 2021
Later among the works it cites.
Learning from eXtreme Bandit Feedback
R. Lopez, I. S. Dhillon, and M. I. Jordan. 2021 · 2021
Later among the works it cites.
Towards Unified Metrics for Accuracy and Diversity for Recommender Systems. In Proc. of the 15th ACM Conference on Recommender Systems (RecSys ’21) . ACM, 75–84
J. Parapar and F. Radlinski. 2021 · 2021
Later among the works it cites.
Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation. In Proc of. the Neural Information Processing Systems Track on Datasets and Benchmarks , Vol. 1
Y. Saito, S. Aihara, M. Matsutani, and Y. Narita. 2021a · 2021
Later among the works it cites.
Counterfactual Learning and Evaluation for Recommender Systems: Foundations, Implementations, and Recent Advances. In Proc. of the 15th ACM Conference on Recommender Systems (RecSys ’21) . ACM, 828–830
Y. Saito and T. Joachims. 2021 · 2021
Later among the works it cites.
Evaluating the Robustness of Click Models to Policy Distributional Shift
R. Deffayet, J. Renders, and M. de Rijke. 2022 · 2022
Later among the works it cites.
Rax: Composable Learning-to-Rank Using JAX. In Proc. of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’22) . ACM, 3051–3060
R. Jagerman, X. Wang, H. Zhuang, Z. Qin, M. Bendersky, and M. Najork. 2022 · 2022
Later among the works it cites.
Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model. In Proc of. the Fifteenth ACM International Conference on Web Search and Data Mining (WSDM ’22) . ACM, 487–497
H. Kiyohara, Y. Saito, T. Matsuhiro, Y. Narita, N. Shimizu, and Y. Yamamoto. 2022 · 2022
Later among the works it cites.
A/B Testing Intuition Busters: Common Misunderstandings in Online Controlled Experiments. In Proc. of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’22) . ACM, 3168–3177
R. Kohavi, A. Deng, and L. Vermeer. 2022 · 2022
Later among the works it cites.
RecPack: An(Other) Experimentation Toolkit for Top-N Recommendation Using Implicit Feedback Data. In Proc. of the 16th ACM Conference on Recommender Systems (RecSys ’22) . ACM, 648–651
L. Michiels, R. Verachtert, and B. Goethals. 2022 · 2022
Later among the works it cites.
Learning-to-Rank at the Speed of Sampling: Plackett-Luce Gradient Estimation with Minimal Computational Complexity. In Proc of. the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22) . ACM, 2266–2271
H. Oosterhuis. 2022 · 2022
Later among the works it cites.
Revisiting the Performance of IALS on Item Recommendation Benchmarks. In Proc. of the 16th ACM Conference on Recommender Systems (RecSys ’22) . ACM, 427–435
S. Rendle, W. Krichene, L. Zhang, and Y. Koren. 2022 · 2022
Later among the works it cites.
Off-Policy Evaluation for Large Action Spaces via Embeddings. In Proc. of the 39th International Conference on Machine Learning (ICML ’22, Vol. 162) . PMLR, 19089–19122
Y. Saito and T. Joachims. 2022 · 2022
Later among the works it cites.
Evaluating Recommender Systems: Survey and Framework
E. Zangerle and C. Bauer. 2022 · 2022
Later among the works it cites.
A Systematic Study on Reproducibility of Reinforcement Learning in Recommendation Systems
E. Cavenaghi, G. Sottocornola, F. Stella, and M. Zanker. 2023 · 2023
Closest in time.
Offline Evaluation for Reinforcement Learning-Based Recommendation: A Critical Issue and Some Alternatives
R. Deffayet, T. Thonet, J. M. Renders, and M. de Rijke. 2023 · 2023
Closest in time.
Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization. In Proc. of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23) . ACM, 249–258
S. Gupta, H. Oosterhuis, and M. de Rijke. 2023 · 2023
Closest in time.
M. Jakimov, A. Buchholz, Y. Stein, and T. Joachims. 2023 · 2023
Closest in time.
A Common Misassumption in Online Experiments with Machine Learning Models
O. Jeunen. 2023a · 2023
Closest in time.
Pessimistic Decision-Making for Recommender Systems
O. Jeunen and B. Goethals. 2023 · 2023
Closest in time.
O. Jeunen and B. London. 2023 · 2023
Closest in time.
A Critical Study on Data Leakage in Recommender System Offline Evaluation
Y. Ji, A. Sun, J. Zhang, and C. Li. 2023 · 2023
Closest in time.
Statistical Challenges in Online Controlled Experiments: A Review of A/B Testing Methodology
N. Larsen, J. Stallrich, S. Sengupta, A. Deng, R. Kohavi, and N. T. Stevens. 2023 · 2023
Closest in time.
Which Tricks are Important for Learning to Rank?. In Proc of. the 40th International Conference on Machine Learning (ICML ’23’, Vol. 202) . PMLR, 23264–23278
I. Lyzhin, A. Ustimenko, A. Gulin, and L. Prokhorenkova. 2023 · 2023
Closest in time.
Doubly Robust Estimation for Correcting Position Bias in Click Feedback for Unbiased Learning to Rank
H. Oosterhuis. 2023 · 2023
Closest in time.
Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling. In Proc. of the 40th International Conference on Machine Learning (ICML ’23, Vol. 202) . PMLR, 29734–29759
Y. Saito, Q. Ren, and T. Joachims. 2023 · 2023
Closest in time.
Variance-Minimizing Augmentation Logging for Counterfactual Evaluation in Contextual Bandits. In Proc of. the Sixteenth ACM International Conference on Web Search and Data Mining (WSDM ’23) . ACM, 967–975
A. D. Tucker and T. Joachims. 2023 · 2023
Closest in time.
Policy-Adaptive Estimator Selection for Off-Policy Evaluation
T. Udagawa, H. Kiyohara, Y. Narita, Y. Saito, and K. Tateno. 2023 · 2023
Closest in time.