Fetching the paper…
Reading the bibliography…
This report documents the program and the outcomes of Dagstuhl Seminar 23031 ``Frontiers of Information Access Experimentation for Research and Education'', which brought together 37 participants from 12 countries.
Über den anschaulichen inhalt der quantentheoretischen kinematik und mechanik
W. Heisenberg · 1927
Earlier work this paper cites.
Experimental and quasi-experimental designs for research
Donald T. Campbell and Julian C. Stanley · 1963
Earlier work this paper cites.
Scale development: theory and applications
Robert F. DeVellis · 1991
Earlier work this paper cites.
When more pain is preferred to less: Adding a better end
Daniel Kahneman, Barbara L. Fredrickson, Charles A. Schreiber, and Donald A. Redelmeier · 1993
Earlier work this paper cites.
User modeling: Recent work, prospects and hazards
Alfred Kobsa · 1993
Earlier work this paper cites.
The philosophy of information retrieval evaluation
Ellen M. Voorhees · 2001
Earlier work this paper cites.
On the recommending of citations for research papers
Sean M. McNee, István Albert, Dan Cosley, Prateep Gopalkrishnan, Shyong K. Lam, Al Mamunur Rashid, Joseph A. Konstan, and John Riedl · 2002
Earlier work this paper cites.
Experimental and quasi-experimental designs for generalized causal inference
William R Shadish, Thomas D Cook, and Donald T. Campbell · 2002
Earlier work this paper cites.
TREC. Experiment and Evaluation in Information Retrieval
D. K. Harman and E. M. Voorhees, editors · 2005
Earlier work this paper cites.
TREC: Experiment and evaluation in information retrieval
E. Voorhees, D.K. Harman, National Institute of Standards, and Technology (US) · 2005
Earlier work this paper cites.
TREC. Experiment and Evaluation in Information Retrieval
D. K. Harman and E. M. Voorhees, editors · 2005
Earlier work this paper cites.
Constructing grounded theory : a practical guide through qualitative analysis
Kathy Charmaz · 2006
Earlier work this paper cites.
Incentives in web studies: Methodological issues and a review
Anja S Göritz · 2006
Earlier work this paper cites.
Reliable information retrieval evaluation with incomplete and biased judgements
Stefan Büttcher, Charles L. A. Clarke, Peter C. K. Yeung, and Ian Soboroff · 2007
Earlier work this paper cites.
An experimental comparison of click position-bias models
Nick Craswell, Onno Zoeter, Michael J. Taylor, and Bill Ramsey · 2008
Earlier work this paper cites.
Can we get rid of trec assessors? using mechanical turk for relevance assessment
Omar Alonso and Stefano Mizzaro · 2009
Earlier work this paper cites.
Methods for evaluating interactive information retrieval systems with users
Diane Kelly · 2009
Earlier work this paper cites.
User-centric evaluation framework for multimedia recommender systems
Bart P. Knijnenburg, Lydia Meesters, Paul Marrow, and Don Bouwhuis · 2009
Earlier work this paper cites.
Improvements that don’t add up: ad-hoc retrieval results since 1998
Timothy G. Armstrong, Alistair Moffat, William Webber, and Justin Zobel · 2009
Earlier work this paper cites.
Most people are not WEIRD
J. Henrich, S. Heine, and A. Norenzayan · 2010
Earlier work this paper cites.
Developing a test collection for the evaluation of integrated search
Marianne Lykke, Birger Larsen, Haakon Lund, and Peter Ingwersen · 2010
Earlier work this paper cites.
Beyond DCG: user behavior as a predictor of a successful search
Ahmed Hassan Awadallah, Rosie Jones, and Kristina Lisa Klinkner · 2010
Earlier work this paper cites.
Do user preferences and evaluation measures line up?
Mark Sanderson, Monica Lestari Paramita, Paul D. Clough, and Evangelos Kanoulas · 2010
Earlier work this paper cites.
Contextual design
Karen Holtzblatt and Hugh R. Beyer · 2011
Earlier work this paper cites.
Information retrieval evaluation
Donna Harman · 2011
Earlier work this paper cites.
Rethinking the recommender research ecosystem: reproducibility, openness, and lenskit
Michael D. Ekstrand, Michael Ludwig, Joseph A. Konstan, and John Riedl · 2011
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Explaining the user experience of recommender systems
Bart P. Knijnenburg, Martijn C. Willemsen, Zeno Gantner, Hakan Soncu, and Chris Newell · 2012
Earlier work this paper cites.
Power analysis for intensive longitudinal studies
Niall Bolger, Gertraud Stadler, and Jean-Philippe Laurenceau · 2012
Earlier work this paper cites.
Explaining the user experience of recommender systems
Bart P. Knijnenburg, Martijn C. Willemsen, Zeno Gantner, Hakan Soncu, and Chris Newell · 2012
Earlier work this paper cites.
Autoregressive and cross-lagged panel analysis for longitudinal data
James P. Selig and Todd D. Little · 2012
Earlier work this paper cites.
A checklist for testing measurement invariance
Rens van de Schoot, Peter Lugtig, and Joop Hox · 2012
Earlier work this paper cites.
Meta-analysis: pitfalls and hints
T. Greco, A. Zangrillo, G. Biondi-Zoccai, and G. Landoni · 2013
Earlier work this paper cites.
The plista dataset
Benjamin Kille, Frank Hopfgartner, Torben Brodt, and Tobias Heintz · 2013
Earlier work this paper cites.
The hawthorne effect and energy awareness
D. Schwartz, B. Fischhoff, T. Krishnamurti, and F. Sowell · 2013
Earlier work this paper cites.
Workshop and challenge on news recommender systems
Mozhgan Tavakolifard, Jon Atle Gulla, Kevin C. Almeroth, Frank Hopfgartner, Benjamin Kille, Till Plumbaum, Andreas Lommatzsch, Torben Brodt, Arthur Bucko, and Tobias Heintz · 2013
Earlier work this paper cites.
Toward identification and adoption of best practices in algorithmic recommender systems research
Joseph A. Konstan and Gediminas Adomavicius · 2013
Earlier work this paper cites.
User perception of differences in recommender algorithms
Michael D. Ekstrand, F. Maxwell Harper, Martijn C. Willemsen, and Joseph A. Konstan · 2014
Earlier work this paper cites.
Measuring surprise in recommender systems
Marius Kaminskas · 2014
Earlier work this paper cites.
Facing the cold start problem in recommender systems
Blerina Lika, Kostas Kolomvatsos, and Stathes Hadjiefthymiades · 2014
Earlier work this paper cites.
An extended data model format for composite recommendation
Alan Said, Babak Loni, Roberto Turrin, and Andreas Lommatzsch · 2014
Earlier work this paper cites.
The impact of search engine selection and sorting criteria on vaccination beliefs and attitudes: Two experiments manipulating google output
Ahmed Allam, Peter Johannes Schulz, and Kent Nakamoto · 2014
Earlier work this paper cites.
Towards a formal framework for utility-oriented measurements of retrieval effectiveness
Marco Ferrante, Nicola Ferro, and Maria Maistro · 2015
Earlier work this paper cites.
Replicable evaluation of recommender systems
Alan Said and Alejandro Bellogín · 2015
Earlier work this paper cites.
A comparison of offline evaluations, online evaluations, and user studies in the context of research-paper recommender systems
Jöran Beel and Stefan Langer · 2015
Earlier work this paper cites.
The search engine manipulation effect (SEME) and its possible impact on the outcomes of elections
Robert Epstein and Ronald E. Robertson · 2015
Earlier work this paper cites.
Reproducibility of data-oriented experiments in e-science (dagstuhl seminar 16041)
Juliana Freire, Norbert Fuhr, and Andreas Rauber · 2016
Earlier work this paper cites.
Understanding the role of latent feature diversification on choice difficulty and satisfaction
Martijn C. Willemsen, Mark P. Graus, and Bart P. Knijnenburg · 2016
Earlier work this paper cites.
Where have all the “workers” gone? a critical analysis of the unrepresentativeness of our samples relative to the labor market in the industrial–organizational psychology literature
Mindy E Bergman and Vanessa A Jean · 2016
Earlier work this paper cites.
The movielens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2016
Earlier work this paper cites.
Recommendations with a purpose
Dietmar Jannach and Gediminas Adomavicius · 2016
Earlier work this paper cites.
A short history of the recsys challenge
Alan Said · 2016
Earlier work this paper cites.
Reproducibility of Data-Oriented Experiments in e-Science (Dagstuhl Seminar 16041)
2016
Earlier work this paper cites.
The netflix recommender system: Algorithms, business value, and innovation
Carlos Alberto Gomez-Uribe and Neil Hunt · 2016
Earlier work this paper cites.
So you’re a program committee member now: On excellence in reviews and meta-reviews and championing submitted work that has merit, 2016
Ken Hinckley · 2016
Cited alongside, same era.
When does relevance mean usefulness and user satisfaction in web search?
Jiaxin Mao, Yiqun Liu, Ke Zhou, Jian-Yun Nie, Jingtao Song, Min Zhang, Shaoping Ma, Jiashen Sun, and Hengliang Luo · 2016
Cited alongside, same era.
Interactions with Search Systems
Ryen W. White · 2016
Cited alongside, same era.
Can Results-Free Review Reduce Publication Bias? The Results and Implications of a Pilot Study
Michael G. Findley, Nathan M. Jensen, Edmund J. Malesky, and Thomas B. Pepinsky · 2016
Cited alongside, same era.
Overview of TREC opensearch 2017
Rolf Jagerman, Krisztian Balog, Philipp Schaer, Johann Schaible, Narges Tavakolpoursaleh, and Maarten de Rijke · 2017
Cited alongside, same era.
Inherent trade-offs in the fair determination of risk scores
Coopetition in IR research
Ellen M. Voorhees · 2020
Later among the works it cites.
Capreolus: A toolkit for end-to-end neural ad hoc retrieval
Andrew Yates, Siddhant Arora, Xinyu Zhang, Wei Yang, Kevin Martin Jose, and Jimmy Lin · 2020
Later among the works it cites.
Models versus satisfaction: Towards a better understanding of evaluation metrics
Fan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie, Weizhi Ma, Min Zhang, and Shaoping Ma · 2020
Later among the works it cites.
Introduction to the special series on results-blind peer review: An experimental analysis on editorial recommendations and manuscript evaluations
Daniel M. Maggin, Rachel E. Robertson, and Bryan G. Cook · 2020
Later among the works it cites.
A living lab architecture for reproducible shared task experimentation
Timo Breuer and Philipp Schaer · 2021
Later among the works it cites.
Explaining recommender systems fairness and accuracy through the lens of data characteristics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jon M. Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2017
Cited alongside, same era.
Searching the enterprise
Udo Kruschwitz and Charlie Hull · 2017
Cited alongside, same era.
Process-tracing methods in decision making: on growing up in the 70s
M. Schulte-Mecklenbeck, J.G. Johnson, U. Böckenholt, D.G. Goldstein, J.E. Russo, N.J. Sullivan, and M.C. Willemsen · 2017
Cited alongside, same era.
Three key affordances for serendipity: Toward a framework connecting environmental and personal factors in serendipitous encounters
Lennart Björneborn · 2017
Cited alongside, same era.
The role of mturk in education research: Advantages, issues, and future directions
D Jake Follmer, Rayne A Sperling, and Hoi K Suen · 2017
Cited alongside, same era.
Beyond the turk: Alternative platforms for crowdsourcing behavioral research
Eyal Peer, Laura Brandimarte, Sonam Samat, and Alessandro Acquisti · 2017
Cited alongside, same era.
Process-tracing methods in decision making: on growing up in the 70s
M. Schulte-Mecklenbeck, J.G. Johnson, U. Böckenholt, D.G. Goldstein, J.E. Russo, N.J. Sullivan, and M.C. Willemsen · 2017
Cited alongside, same era.
Yashar Deldjoo, Alejandro Bellogín, and Tommaso Di Noia · 2021
Later among the works it cites.
Managing bias in human-annotated data: Moving beyond bias removal
Gianluca Demartini, Kevin Roitero, and Stefano Mizzaro · 2021
Later among the works it cites.
Principled multi-aspect evaluation measures of rankings
Maria Maistro, Lucas Chaves Lima, Jakob Grue Simonsen, and Christina Lioma · 2021
Later among the works it cites.
Evaluating Information Retrieval and Access Tasks – NTCIR’s Legacy of Research Impact
T. Sakai, D. W. Oard, and N. Kando, editors · 2021
Later among the works it cites.
EXAM: how to evaluate retrieve-and-generate systems for users who do not (yet) know what they want
David P. Sander and Laura Dietz · 2021
Later among the works it cites.
Overview of lilas 2021 - living labs for academic search
Philipp Schaer, Timo Breuer, Leyla Jael Castro, Benjamin Wolff, Johann Schaible, and Narges Tavakolpoursaleh · 2021
Later among the works it cites.
Recsys 2021 challenge workshop: Fairness-aware engagement prediction at scale on twitter’s home timeline
Vito Walter Anelli, Saikishore Kalloori, Bruce Ferwerda, Luca Belli, Alykhan Tejani, Frank Portman, Alexandre Lung-Yut-Fong, Ben Chamberlain, Yuanpu Xie, Jonathan Hunt, Michael M. Bronstein, and Wenzhe Shi · 2021
Later among the works it cites.
Improving accountability in recommender systems research through reproducibility
Alejandro Bellogín and Alan Said · 2021
Later among the works it cites.
Data collection via online platforms: Challenges and recommendations for future research
Alexander Newman, Yuen Lam Bavik, Matthew Mount, and Bo Shao · 2021
Later among the works it cites.
Overview of lilas 2021 - living labs for academic search (extended overview)
Philipp Schaer, Timo Breuer, Leyla Jael Castro, Benjamin Wolff, Johann Schaible, and Narges Tavakolpoursaleh · 2021
Later among the works it cites.
Overview of the CLEF ehealth evaluation lab 2021
Hanna Suominen, Lorraine Goeuriot, Liadh Kelly, Laura Alonso Alemany, Elias Bassani, Nicola Brew-Sam, Viviana Cotik, Darío Filippo, Gabriela González Sáez, Franco Luque, Philippe Mulhem, Gabriella Pasi, Roland Roller, Sandaru Seneviratne, Rishabh Upadhyay, Jorge Vivaldi, Marco Viviani, and Chenchen Xu · 2021
Later among the works it cites.
Knowledge task survey
Elaine Toms, Sophia Althammer, Allan Hanbury, Wojciech Kusa, Ginar Santika Niwanputri, Ian Ruthven, Ayah Soufan, and Vasileios Stamatis · 2021
Later among the works it cites.
Assessing top-k preferences
Charles L. A. Clarke, Alexandra Vtyurina, and Mark D. Smucker · 2021
Later among the works it cites.
EXAM: how to evaluate retrieve-and-generate systems for users who do not (yet) know what they want
David P. Sander and Laura Dietz · 2021
Later among the works it cites.
Overview of touché 2021: Argument retrieval
Alexander Bondarenko, Lukas Gienapp, Maik Fröbe, Meriem Beloucif, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, and Matthias Hagen · 2021
Later among the works it cites.
Assessing viewpoint diversity in search results using ranking fairness metrics
Tim Draws, Nava Tintarev, and Ujwal Gadiraju · 2021
Later among the works it cites.
This is not what we ordered: Exploring why biased search result rankings affect user attitudes on debated topics
Tim Draws, Nava Tintarev, Ujwal Gadiraju, Alessandro Bozzon, and Benjamin Timmermans · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Sasko, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander M. Rush, and Thomas Wolf · 2021
Later among the works it cites.
Significant improvements over the state of the art? A case study of the MS MARCO document ranking leaderboard
Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, and Emine Yilmaz · 2021
Later among the works it cites.
Simplified data wrangling with ir_datasets
Sean MacAvaney, Andrew Yates, Sergey Feldman, Doug Downey, Arman Cohan, and Nazli Goharian · 2021
Later among the works it cites.
Pyterrier: Declarative experimentation in python from BM25 to dense retrieval
Craig Macdonald, Nicola Tonellotto, Sean MacAvaney, and Iadh Ounis · 2021
Later among the works it cites.
How to write CHI papers, online edition
Lennart E. Nacke · 2021
Later among the works it cites.
The jasp guidelines for conducting and reporting a bayesian analysis
Johnny van Doorn, Don van den Bergh, Udo Böhm, Fabian Dablander, Koen Derks, Tim Draws, Alexander Etz, Nathan J Evans, Quentin F Gronau, Julia M Haaf, et al · 2021
Later among the works it cites.
ir_metadata: An extensible metadata schema for IR experiments
Timo Breuer, Jüri Keller, and Philipp Schaer · 2022
Later among the works it cites.
Fairness in information access systems
Michael D. Ekstrand, Anubrata Das, Robin Burke, and Fernando Diaz · 2022
Later among the works it cites.
Socio-economic diversity in human annotations
Shaoyang Fan, Pinar Barlas, Evgenia Christoforou, Jahna Otterbacher, Shazia W. Sadiq, and Gianluca Demartini · 2022
Later among the works it cites.
How does the crowd impact the model? A tool for raising awareness of social bias in crowdsourced training data
Periklis Perikleous, Andreas Kafkalias, Zenonas Theodosiou, Pinar Barlas, Evgenia Christoforou, Jahna Otterbacher, Gianluca Demartini, and Andreas Lanitis · 2022
Later among the works it cites.
A survey on the fairness of recommender systems
Yifan Wang, Weizhi Ma, Min Zhang, Yiqun Liu, and Shaoping Ma · 2022
Later among the works it cites.
When measurement misleads: The limits of batch assessment of retrieval systems
J. Zobel · 2022
Later among the works it cites.
Experimental IR Meets Multilinguality, Multimodality, and Interaction - 13th International Conference of the CLEF Association, CLEF 2022, Bologna, Italy, September 5-8, 2022, Proceedings
Alberto Barrón-Cedeño, Giovanni Da San Martino, Mirko Degli Esposti, Fabrizio Sebastiani, Craig Macdonald, Gabriella Pasi, Allan Hanbury, Martin Potthast, Guglielmo Faggioli, and Nicola Ferro, editors · 2022
Later among the works it cites.
ir_metadata: An extensible metadata schema for IR experiments
Timo Breuer, Jüri Keller, and Philipp Schaer · 2022
Later among the works it cites.
Kinds of replication: Examining the meanings of “conceptual replication” and “direct replication”
Maarten Derksen and Jill Morawski · 2022
Later among the works it cites.
From reality to world. a critical perspective on AI fairness
Jean-Marie John-Mathews, Dominique Cardon, and Christine Balagué · 2022
Later among the works it cites.
Recsys challenge 2022: Fashion purchase prediction
Nick Landia, Frederick Cheung, Donna North, Saikishore Kalloori, Abhishek Srivastava, and Bruce Ferwerda · 2022
Later among the works it cites.
Exploring the longitudinal effects of nudging on users’ music genre exploration behavior and listening preferences
Yu Liang and Martijn C. Willemsen · 2022
Later among the works it cites.
Longitudinal user experience studies in the iot domain: a brief panorama and challenges to overcome
Bianca Melo, Rossana M. de Castro Andrade, and Ticianne Darin · 2022
Later among the works it cites.
Overview of touché 2022: Argument retrieval
Alexander Bondarenko, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Gurcke, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, and Matthias Hagen · 2022
Later among the works it cites.
Emerging trends: Sota-chasing
Kenneth Ward Church and Valia Kordoni · 2022
Later among the works it cites.
How do you test a test?: A multifaceted examination of significance tests
Nicola Ferro and Mark Sanderson · 2022
Later among the works it cites.
The role of historical and contextual knowledge in enterprise search
Marianne Lykke, Ann Bygholm, Louise Bak Søndergaard, and Katriina Byström · 2022
Later among the works it cites.
When measurement misleads: The limits of batch assessment of retrieval systems
J. Zobel · 2022
Later among the works it cites.
Experimental standards for deep learning in natural language processing research, 2022
Dennis Ulmer, Elisa Bassignana, Max Müller-Eberstein, Daniel Varab, Mike Zhang, Rob van der Goot, Christian Hardmeier, and Barbara Plank · 2022
Later among the works it cites.
A unifying and general account of fairness measurement in recommender systems
Enrique Amigó, Yashar Deldjoo, Stefano Mizzaro, and Alejandro Bellogín · 2023
Closest in time.
On the role of human and machine metadata in relevance judgment tasks
Jiechen Xu, Lei Han, Shazia Sadiq, and Gianluca Demartini · 2023
Closest in time.
Evaluating recommender systems: Survey and framework
Eva Zangerle and Christine Bauer · 2023
Closest in time.
Viewpoint diversity in search results
Tim Draws, Nirmal Roy, Oana Inel, Alisa Rieger, Rishav Hada, Mehmet Orcun Yalcin, Benjamin Timmermans, and Nava Tintarev · 2023
Closest in time.
Shared Tasks as Tutorials: A Methodical Approach
Theresa Elstner, Frank Loebe, Yamen Ajjour, Christopher Akiki, Alexander Bondarenko, Maik Fröbe, Lukas Gienapp, Nikolay Kolyada, Janis Mohr, Stephan Sandfuchs, Matti Wiegmann, Jörg Frochte, Nicola Ferro, Sven Hofmann, Benno Stein, Matthias Hagen, and Martin Potthast · 2023
Closest in time.
Continuous Integration for Reproducible Shared Tasks with TIRA.io
Maik Fröbe, Matti Wiegmann, Nikolay Kolyada, Bastian Grahm, Theresa Elstner, Frank Loebe, Matthias Hagen, Benno Stein, and Martin Potthast · 2023
Closest in time.
Evaluating recommender systems: Survey and framework
Eva Zangerle and Christine Bauer · 2023
Closest in time.