Fetching the paper…
Reading the bibliography…
The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR) in zero-shot settings, garnering considerable attention within the research community.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 1906
Earlier work this paper cites.
Www’18 open challenge: Financial opinion mining and question answering
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018 · 1942
Earlier work this paper cites.
Why inverse document frequency?
Kishore Papineni. 2001 · 2001
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2004
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over BERT
Omar Khattab and Matei Zaharia. 2020 · 2004
Earlier work this paper cites.
KLEJ: comprehensive benchmark for polish language understanding
Piotr Rybak, Robert Mroczkowski, Janusz Tracz, and Ireneusz Gawlik. 2020 · 2005
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
National corpus of polish
Adam Przepiórkowski, Mirosław Bańko, Rafał L Górski, Barbara Lewandowska-Tomaszczyk, Marek Łaziński, and Piotr Pęzik. 2011 · 2011
Earlier work this paper cites.
CLIMATE-FEVER: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan L. Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020 · 2012
Earlier work this paper cites.
Cqadupstack: A benchmark data set for community question-answering research
Doris Hoogeveen, Karin M. Verspoor, and Timothy Baldwin. 2015 · 2015
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
Cqadupstack : Gold or silver ?
Doris Hoogeveen, Karin M. Verspoor, and Timothy Baldwin. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Dbpedia-entity v2: A test collection for entity search
Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017 · 2017
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
Retrieval of the best counterargument without prior topic knowledge
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Polish corpus of wrocław university of technology 1.3
Marcin Oleksy, Michał Marcińczuk, Marek Maziarz, Tomasz Bernaś, Jan Wieczorek, Agnieszka Turek, Dominika Fikus, Michał Wolski, Marek Pustowaruk, Jan Kocoń, et al. 2019 · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Trec-covid: Constructing a pandemic information retrieval test collection
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021 · 2021
Later among the works it cites.
Mr. tydi: A multi-lingual benchmark for dense retrieval
Xinyu Zhang, Xueguang Ma, Peng Shi, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
This is the way: designing and compiling lepiszcze, a comprehensive nlp benchmark for polish
Łukasz Augustyniak, Kamil Tagowski, Albert Sawczyn, Denis Janiak, Roman Bartusiak, Adrian Szymczak, Marcin Wątroba, Arkadiusz Janz, Piotr Szymański, Mikołaj Morzy, Tomasz Kajdanowicz, and Maciej Piasecki. 2022 · 2022
Later among the works it cites.
Inpars: Unsupervised dataset generation for information retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Overview of touché 2020: Argument retrieval: Extended abstract
Alexander Bondarenko, Maik Fröbe, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, and Matthias Hagen. 2020 · 2020
Cited alongside, same era.
Specter: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020 · 2020
Cited alongside, same era.
Pre-training polish transformer-based language models at scale
Sławomir Dadas, Michał Perełkiewicz, and Rafał Poświata. 2020 · 2020
Cited alongside, same era.
Document ranking with a pretrained sequence-to-sequence model
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
OPUS-MT – building open translation services for the world
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Cited alongside, same era.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Cited alongside, same era.
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Later among the works it cites.
Evaluation of transfer learning for polish with a text-to-text model
Aleksandra Chrabrowa, Łukasz Dragan, Karol Grzegorczyk, Dariusz Kajtoch, Mikołaj Koszowski, Robert Mroczkowski, and Piotr Rybak. 2022 · 2022
Later among the works it cites.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Later among the works it cites.
Aggretriever: A simple approach to aggregate textual representation for robust dense passage retrieval
Sheng-Chieh Lin, Minghan Li, and Jimmy Lin. 2022 · 2022
Later among the works it cites.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022 · 2022
Later among the works it cites.
LaPraDoR: Unsupervised pretrained dense retriever for zero-shot text retrieval
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2022 · 2022
Later among the works it cites.
Poleval 2022/23 challenge tasks and results
Łukasz Kobyliński, Maciej Ogrodniczuk, Piotr Rybak, Piotr Przybyła, Piotr Pęzik, Agnieszka Mikołajczyk, Wojciech Janowski, Michał Marcińczuk, and Aleksander Smywiński-Pohl. 2023 · 2022
Later among the works it cites.
Hybrid retrievers with generative re-rankers
Marek Kozłowski. 2023 · 2023
Closest in time.
Passage retrieval of polish texts using okapi bm25 and an ensemble of cross encoders
Jakub Pokrywka. 2023 · 2023
Closest in time.
MAUPQA: Massive automatically-created Polish question answering dataset
Piotr Rybak. 2023 · 2023
Closest in time.
Multi-index retrieve and rerank with sequence-to-sequence model
Konrad Wojtasik. 2023 · 2023
Closest in time.