Fetching the paper…
Reading the bibliography…
In this chapter, we consider generative information retrieval evaluation from two distinct but interrelated perspectives.
Rajput, S., Pavlu, V., Golbus, P. B. and Aslam, J. A. [2011], A nugget-based test collection construction paradigm, in
1948
Earlier work this paper cites.
Jardine, N. and van Rijsbergen, C. J. [1971], ‘The use of hierarchic clustering in information retrieval’, Inf. Storage Retr
1971
Earlier work this paper cites.
Spärck Jones, K. and Van Rijsbergen, C. [1975], Report on the need for and provision of an ‘ideal’ information retrieval test collection, Technical Report 5266, British Library Research and Development Report. https://cir.nii.ac.jp/crid/1570572699089480448
1975
Earlier work this paper cites.
Robertson, S. [1977], ‘The probability ranking principle in IR’, Journal of Documentation
1977
Earlier work this paper cites.
Croft, W. B. and Harper, D. J. [1979], ‘Using probabilistic models of document retrieval without relevance information’, Journal of Documentation
1979
Earlier work this paper cites.
van Rijsbergen, C. [1979], Information retrieval
1979
Earlier work this paper cites.
Harman, D. [1992], ‘User-friendly systems instead of user-friendly front-ends’, Journal of the American Society for Information Science (JASIST)
1992
Earlier work this paper cites.
Hearst, M. A. and Pedersen, J. O. [1996], Reexamining the cluster hypothesis: Scatter/gather on retrieval results, in
1996
Earlier work this paper cites.
Tombros, A. and Sanderson, M. [1998], Advantages of query biased summaries in information retrieval, in
1998
Earlier work this paper cites.
Zobel, J. [1998], How reliable are the results of large-scale information retrieval experiments?, in
1998
Earlier work this paper cites.
Järvelin, K. and Kekäläinen, J. [2002], ‘Cumulated gain-based evaluation of IR techniques’, ACM Transactions on Information Systems (TOIS)
2002
Earlier work this paper cites.
Bernstein, Y. and Zobel, J. [2005], Redundant documents and search effectiveness, in
2005
Earlier work this paper cites.
Büttcher, S. and Clarke, C. L. A. [2005], Efficiency vs. effectiveness in terabyte-scale information retrieval, in
2005
Earlier work this paper cites.
Voorhees, E., Harman, D., of Standards, N. I. and (US), T. [2005], TREC: Experiment and evaluation in information retrieval
2005
Earlier work this paper cites.
Buckley, C., Dimmick, D., Soboroff, I. and Voorhees, E. [2006], Bias and the limits of pooling, in
2006
Earlier work this paper cites.
Rieh, S. Y. and Xie, H. I. [2006], ‘Analysis of multiple query reformulations on the eeb: The interactive information retrieval context’, Information Processing & Management (IP&M)
2006
Earlier work this paper cites.
Turpin, A. and Scholer, F. [2006], User performance versus precision measures for simple search tasks, in
2006
Earlier work this paper cites.
Church, K., Smyth, B., Cotter, P. and Bradley, K. [2007], ‘Mobile information access: A study of emerging search behavior on the mobile internet’, ACM Transactions on the Web (TWEB)
2007
Earlier work this paper cites.
Clarke, C. L. A., Agichtein, E., Dumais, S. T. and White, R. W. [2007], The influence of caption features on clickthrough patterns in web search, in
2007
Earlier work this paper cites.
Glover, E. [2007], The real world web search problem: bridging the gap between academic and commercial understanding of issues and methods, in
2007
Earlier work this paper cites.
Bailey, P., Craswell, N., Soboroff, I., Thomas, P., de Vries, A. P. and Yilmaz, E. [2008], Relevance assessment: are judges exchangeable and does it matter, in
2008
Earlier work this paper cites.
Cao, H., Jiang, D., Pei, J., He, Q., Liao, Z., Chen, E. and Li, H. [2008], Context-aware query suggestion by mining click- through and session data, in
2008
Earlier work this paper cites.
Craswell, N., Zoeter, O., Taylor, M. J. and Ramsey, B. [2008], An experimental comparison of click position-bias models, in
2008
Earlier work this paper cites.
Moffat, A. and Zobel, J. [2008], ‘Rank-biased precision for measurement of retrieval effectiveness’, ACM Transactions on Information Systems (TOIS)
2008
Earlier work this paper cites.
Chakrabarti, D., Kumar, R. and Punera, K. [2009], Quicklink selection for navigational query results, in
2009
Earlier work this paper cites.
Chapelle, O., Metlzer, D., Zhang, Y. and Grinspan, P. [2009], Expected reciprocal rank for graded relevance, in
2009
Earlier work this paper cites.
White, R. W., Dumais, S. T. and Teevan, J. [2009], Characterizing the influence of domain expertise on web search behavior, in
2009
Earlier work this paper cites.
Al-Maskari, A. and Sanderson, M. [2010], ‘A review of factors influencing user satisfaction in information retrieval’, Journal of the American Society for Information Science (JASIST)
2010
Earlier work this paper cites.
Azzopardi, L., Järvelin, K., Kamps, J. and Smucker, M. D. [2010], ‘Report on the SIGIR 2010 workshop on the simulation of interaction’, SIGIR Forum
2010
Earlier work this paper cites.
Bailey, P., Craswell, N., White, R. W., Chen, L., Satyanarayana, A. and Tahaghoghi, S. M. M. [2010], Evaluating whole-page relevance, in
2010
Cited alongside, same era.
Sanderson, M. [2010], ‘Test collection based evaluation of information retrieval systems’, Foundations and Trends® in Information Retrieval
2010
Cited alongside, same era.
Sanderson, M., Scholer, F. and Turpin, A. [2010], Relatively relevant: Assessor shift in document judgements, in
2010
Cited alongside, same era.
Torres, S. D., Hiemstra, D. and Serdyukov, P. [2010], Query log analysis in the context of information retrieval for children, in
2010
Cited alongside, same era.
Haas, K., Mika, P., Tarjan, P. and Blanco, R. [2011], Enhanced results for web search, in
2011
Cited alongside, same era.
Zhang, Y., Hu, C., Liu, Y., Fang, H. and Lin, J. [2021], Learning to rank in the age of Muppets: Effectiveness–efficiency tradeoffs in multi-stage ranking, in
2021
Later among the works it cites.
Alaofi, M., Gallagher, L., McKay, D., Saling, L. L., Sanderson, M., Scholer, F., Spina, D. and White, R. W. [2022], Where do queries come from?, in
2022
Later among the works it cites.
Bonifacio, L. H., Abonizio, H. Q., Fadaee, M. and Nogueira, R. F. [2022], Inpars: Unsupervised dataset generation for information retrieval, in
2022
Later among the works it cites.
Culpepper, J. S., Faggioli, G., Ferro, N. and Kurland, O. [2022], ‘Topic difficulty: Collection and query formulation effects’, ACM Transactions on Information Systems (TOIS)
2022
Later among the works it cites.
Moffat, A., Mackenzie, J., Thomas, P. and Azzopardi, L. [2022], A flexible framework for offline effectiveness metrics, in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scholer, F., Turpin, A. and Sanderson, M. [2011], Quantifying test collection quality based on the consistency of relevance judgements, in
2011
Cited alongside, same era.
Wang, L., Lin, J. and Metzler, D. [2011], A cascade ranking model for efficient ranked retrieval, in
2011
Cited alongside, same era.
Smucker, M. D. and Clarke, C. L. A. [2012], Time-based calibration of effectiveness measures, in
2012
Cited alongside, same era.
Navalpakkam, V., Jentzsch, L., Sayres, R., Ravi, S., Ahmed, A. and Smola, A. J. [2013], Measurement and modeling of eye-mouse behavior in the presence of nonlinear page layouts, in
2013
Cited alongside, same era.
Arapakis, I., Bai, X. and Cambazoglu, B. B. [2014], Impact of response latency on user behavior in web search, in
2014
Cited alongside, same era.
Kurland, O. [2014], The cluster hypothesis in information retrieval, in
2014
Cited alongside, same era.
Maxwell, D. and Azzopardi, L. [2014], Stuck in traffic: how temporal delays affect search behaviour, in
2014
Cited alongside, same era.
2022
Later among the works it cites.
Penha, G., Câmara, A. and Hauff, C. [2022], Evaluating the robustness of retrieval pipelines with query variation generators, in
2022
Later among the works it cites.
Voorhees, E. M., Craswell, N. and Lin, J. [2022], Too many relevants: Whither cranfield test collections?, in
2022
Later among the works it cites.
Alaofi, M., Gallagher, L., Sanderson, M., Scholer, F. and Thomas, P. [2023], Can generative llms create query variants for test collections? an exploratory study, in
2023
Later among the works it cites.
Arabzadeh, N., Kmet, O., Carterette, B., Clarke, C. L. A., Hauff, C. and Chandar, P. [2023], A is for adele: An offline evaluation metric for instant search, in
2023
Later among the works it cites.
Balog, K. and Zhai, C. [2023], ‘User simulation for evaluating information access systems’. https://doi.org/10.48550/arXiv.2306.08550
2023
Later among the works it cites.
Chen, J., Zhang, R., Guo, J., de Rijke, M., Chen, W., Fan, Y. and Cheng, X. [2023], Continual learning for generative retrieval over dynamic corpora, in
2023
Later among the works it cites.
Chen, J., Zhang, R., Guo, J., de Rijke, M., Liu, Y., Fan, Y. and Cheng, X. [2023], A unified generative retriever for knowledge- intensive language tasks via prompt learning, in
2023
Later among the works it cites.
Deckers, N., Fröbe, M., Kiesel, J., Pandolfo, G., Schröder, C., Stein, B. and Potthast, M. [2023], The infinite index: Information retrieval on generative text-to-image models, in
2023
Later among the works it cites.
Faggioli, G., Dietz, L., Clarke, C. L. A., Demartini, G., Hagen, M., Hauff, C., Kando, N., Kanoulas, E., Potthast, M., Stein, B. and Wachsmuth, H. [2023], Perspectives on large language models for relevance judgment, in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Hämäläinen, P., Tavast, M. and Kunnari, A. [2023], Evaluating large language models in generating synthetic HCI research data: a case study, in
2023
Later among the works it cites.
Oliveira, B. and Lopes, C. T. [2023], The evolution of web search user interfaces - an archaeological analysis of google search engine result pages, in
2023
Later among the works it cites.
Penha, G., Palumbo, E., Aziz, M., Wang, A. and Bouchard, H. [2023], Improving content retrievability in search with controllable query generation, in
2023
Later among the works it cites.
Pera, M. S., Murgia, E., Landoni, M., Huibers, T. and Aliannejadi, M. [2023], Where a little change makes a big difference: A preliminary exploration of children’s queries, in
2023
Later among the works it cites.
Pradeep, R., Hui, K., Gupta, J., Lelkes, Á. D., Zhuang, H., Lin, J., Metzler, D. and Tran, V. Q. [2023], How does generative retrieval scale to millions of passages?, in
2023
Later among the works it cites.
Sun, W., Yan, L., Chen, Z., Wang, S., Zhu, H., Ren, P., Chen, Z., Yin, D., de Rijke, M. and Ren, Z. [2023], Learning to tokenize for generative retrieval, in
2023
Later among the works it cites.
2023
Later among the works it cites.
Yang, T., Song, M., Zhang, Z., Huang, H., Deng, W., Sun, F. and Zhang, Q. [2023], Auto search indexer for end-to-end document retrieval, in
2023
Later among the works it cites.
Abbasiantaeb, Z. and Aliannejadi, M. [2024], ‘Generate then retrieve: Conversational response retrieval using llms as answer and query generators’
2024
Closest in time.
Engelmann, B., Breuer, T., Friese, J. I., Schaer, P. and Fuhr, N. [2024], Context-driven interactive query simulations based on generative large language models, in
2024
Closest in time.
2024
Closest in time.
Mackie, I., Chatterjee, S. and Dalton, J. [2023], Generative relevance feedback with large language models, in
2031
Closest in time.