Fetching the paper…
Reading the bibliography…
We carried out a reproducibility study of InPars, which is a method for unsupervised training of neural rankers (Bonifacio et al., 2022).
Understanding inverse document frequency: on theoretical arguments for IDF
Stephen Robertson · 2004
Earlier work this paper cites.
Overview of the trec 2004 robust retrieval track
Ellen Voorhees · 2004
Earlier work this paper cites.
High accuracy retrieval with multiple nested ranker
Irina Matveeva, Chris Burges, Timo Burkard, Andy Laucius, and Leon Wong · 2006
Earlier work this paper cites.
Open-domain question-answering
John M. Prager · 2006
Earlier work this paper cites.
Bias and the limits of pooling for large collections
Chris Buckley, Darrin Dimmick, Ian Soboroff, and Ellen M. Voorhees · 2007
Earlier work this paper cites.
Improving efficient neural ranking models with cross-architecture knowledge distillation, 2020
Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury · 2010
Earlier work this paper cites.
A cascade ranking model for efficient ranked retrieval
Lidan Wang, Jimmy Lin, and Donald Metzler · 2011
Earlier work this paper cites.
Learning small-size DNN with output-distribution-based criteria
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N. Smith · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Synthetic QA corpora generation with roundtrip consistency
Chris Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, and Michael Collins · 2019
Earlier work this paper cites.
Natural Questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Rodrigo Nogueira and Kyunghyun Cho · 2019
Cited alongside, same era.
Flexible retrieval with NMSLIB and FlexNeuART
Leonid Boytsov and Eric Nyberg · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees · 2020
Cited alongside, same era.
Crowd worker strategies in relevance judgment tasks
Lei Han, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, and Gianluca Demartini · 2020
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2021
Later among the works it cites.
Simplified data wrangling with ir_datasets
Sean MacAvaney, Andrew Yates, Sergey Feldman, Doug Downey, Arman Cohan, and Nazli Goharian · 2021
Later among the works it cites.
A systematic evaluation of transfer learning and pseudo-labeling with bert-based ranking models
Iurii Mokrii, Leonid Boytsov, and Pavel Braslavski · 2021
Later among the works it cites.
Large dual encoders are generalizable retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hern’andez ’Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang · 2021
Later among the works it cites.
Synthetic target domain supervision for open retrieval QA
Revanth Gangi Reddy, Bhavani Iyer, Md. Arafat Sultan, Rong Zhang, Avirup Sil, Vittorio Castelli, Radu Florian, and Salim Roukos · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Contrastive representation learning: A framework and review
Phuc H. Le-Khac, Graham Healy, and Alan F. Smeaton · 2020
Cited alongside, same era.
Distilling dense representations for ranking using tightly-coupled teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin · 2020
Cited alongside, same era.
On the stability of fine-tuning BERT: misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow · 2020
Cited alongside, same era.
Document ranking with a pretrained sequence-to-sequence model
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, and William R. Hersh · 2020
Cited alongside, same era.
ERNIE 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Cited alongside, same era.
Later among the works it cites.
Generating datasets with pretrained language models
Timo Schick and Hinrich Schütze · 2021
Later among the works it cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
Inpars: Unsupervised dataset generation for information retrieval
Luiz Henrique Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira · 2022
Later among the works it cites.
Promptagator: Few-shot dense retrieval from 8 examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang · 2022
Later among the works it cites.
Precise zero-shot dense retrieval without relevance labels, 2022
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan · 2022
Later among the works it cites.
Domain adaptation for dense retrieval through self-supervision by pseudo-relevance labeling
Minghan Li and Éric Gaussier · 2022
Later among the works it cites.
In defense of cross-encoders for zero-shot retrieval, 2022
Guilherme Rosa, Luiz Bonifacio, Vitor Jeronymo, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira · 2022
Later among the works it cites.
Improving passage retrieval with zero-shot question generation
Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al · 2022
Later among the works it cites.
GPL: generative pseudo labeling for unsupervised domain adaptation of dense retrieval
Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych · 2022
Later among the works it cites.
Inpars-v2: Large language models as efficient dataset generators for information retrieval, 2023
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira · 2023
Closest in time.