Fetching the paper…
Reading the bibliography…
The information retrieval community has recently witnessed a revolution due to large pretrained transformer models.
Okapi at trec-3
SE ROBERTSON, S WALKER, S JONES, MM HANCOCK-BEAULIEU, and M GATFORD. 1995 · 1995
Earlier work this paper cites.
Overview of the eighth text retrieval conference (trec-8)
Ellen M. Voorhees and Donna K. Harman. 1999 · 1999
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2004
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over BERT
Omar Khattab and Matei Zaharia. 2020 · 2004
Earlier work this paper cites.
Dare: Data augmented relation extraction with gpt-2
Yannis Papanikolaou and Andrea Pierleoni. 2020 · 2004
Earlier work this paper cites.
Overview of the trec 2004 robust retrieval track
Ellen Voorhees. 2005 · 2004
Earlier work this paper cites.
Overview of trec 2004
Ellen M. Voorhees. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
High accuracy retrieval with multiple nested ranker
Irina Matveeva, Chris Burges, Timo Burkard, Andy Laucius, and Leon Wong. 2006 · 2006
Earlier work this paper cites.
Parade: Passage representation aggregation for document reranking
Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun. 2020 · 2008
Earlier work this paper cites.
A cascade ranking model for efficient ranked retrieval
Lidan Wang, Jimmy Lin, and Donald Metzler. 2011 · 2011
Earlier work this paper cites.
The curse of dense low-dimensional information retrieval for large index sizes
Nils Reimers and Iryna Gurevych. 2020 · 2012
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Efficient cost-aware cascade ranking in multi-stage retrieval
Ruey-Cheng Chen, Luke Gallagher, Roi Blanco, and J Shane Culpepper. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
Cascade ranking for operational e-commerce search
Shichen Liu, Fei Xiao, Wenwu Ou, and Luo Si. 2017 · 2017
Earlier work this paper cites.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi. 2018 · 2018
Earlier work this paper cites.
Deeper text understanding for ir with contextual neural language modeling
Zhuyun Dai and Jamie Callan. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Do not have enough data? deep learning to the rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. 2020 · 2020
Cited alongside, same era.
Overview of the TREC 2020 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2020
Cited alongside, same era.
Simulated chats for building dialog systems: Learning to generate conversations from instructions
Biswesh Mohapatra, Gaurav Pandey, Danish Contractor, and Sachindra Joshi. 2021 · 2021
Later among the works it cites.
The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
Co{da}: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding
Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Weizhu Chen, and Jiawei Han. 2021 · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020 · 2020
Cited alongside, same era.
Document ranking with a pretrained sequence-to-sequence model
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020 · 2020
Cited alongside, same era.
H2oloo at trec 2020: When all you got is a hammer… deep learning, health misinformation, and precision medicine
Ronak Pradeep, Xueguang Ma, Xinyu Zhang, H. Cui, Ruizhou Xu, Rodrigo Nogueira, Jimmy J. Lin, and D. Cheriton. 2020a · 2020
Cited alongside, same era.
H2oloo at trec 2020: When all you got is a hammer… deep learning, health misinformation, and precision medicine
Ronak Pradeep, Xueguang Ma, Xinyu Zhang, Hang Cui, Ruizhou Xu, Rodrigo Nogueira, and Jimmy Lin. 2020b · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen Voorhees, Lucy Lu Wang, and William R Hersh. 2020 · 2020
Cited alongside, same era.
G-daug: Generative data augmentation for commonsense reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020 · 2020
Cited alongside, same era.
Ori Ram, Gal Shachaf, Omer Levy, Jonathan Berant, and Amir Globerson. 2021 · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021 · 2021
Later among the works it cites.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021 · 2021
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze. 2021b · 2021
Later among the works it cites.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Gpl: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval
Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Language models are few-shot multilingual learners
Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Adaptsum: Towards low-resource domain adaptation for abstractive summarization
Tiezheng Yu, Zihan Liu, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022 · 2022
Closest in time.
Generating training data with language models: Towards zero-shot language understanding
Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022 · 2022
Closest in time.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022 · 2022
Closest in time.
Selecting parallel in-domain sentences for neural machine translation using monolingual texts
Javad Pourmostafa Roshan Sharami, Dimitar Shterionov, and Pieter Spronck. 2022 · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022 · 2022
Closest in time.