Fetching the paper…
Reading the bibliography…
The MS MARCO ranking dataset has been widely used for training deep learning models for IR tasks, achieving considerable effectiveness on diverse zero-shot scenarios.
Cross-lingual Language Model Pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 1904
Earlier work this paper cites.
On the Cross-lingual Transferability of Monolingual Representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2019 · 1910
Earlier work this paper cites.
Multi-Stage Document Ranking with BERT
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019 · 1910
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 1911
Earlier work this paper cites.
Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering
Casimiro Pio Carrino, Marta R. Costa-jussà, and José A. R. Fonollosa. 2019 · 1912
Earlier work this paper cites.
Teaching a New Dog Old Tricks: Resurrecting Multilingual Retrieval Using Zero-shot Learning
Sean MacAvaney, Luca Soldaini, and Nazli Goharian. 2019 · 1912
Earlier work this paper cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2002
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020 · 2003
Earlier work this paper cites.
Document Ranking with a Pretrained Sequence-to-Sequence Model
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2020 · 2003
Earlier work this paper cites.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab and Matei Zaharia. 2020a · 2004
Earlier work this paper cites.
Bias and the limits of pooling for large collections
Chris Buckley, Darrin Dimmick, Ian Soboroff, and Ellen Voorhees. 2007 · 2007
Earlier work this paper cites.
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2020 · 2007
Earlier work this paper cites.
PARADE: Passage Representation Aggregation for Document Reranking
Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun. 2020 · 2008
Earlier work this paper cites.
A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. 2020 · 2010
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020 · 2010
Cited alongside, same era.
The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes
Nils Reimers and Iryna Gurevych. 2020 · 2012
Cited alongside, same era.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Cited alongside, same era.
Syntactic Data Augmentation Increases Robustness to Inference Heuristics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 2339–2352
Junghyun Min, R. Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen. 2020 · 2020
Later among the works it cites.
H2oloo at TREC 2020: When all you got is a hammer… Deep Learning, Health Misinformation, and Precision Medicine. In TREC
Ronak Pradeep, Xueguang Ma, Xinyu Zhang, H. Cui, Ruizhou Xu, Rodrigo Nogueira, Jimmy J. Lin, and D. Cheriton. 2020 · 2020
Later among the works it cites.
OPUS-MT — Building open translation services for the World. In Proceedings of the 22nd Annual Conferenec of the European Association for Machine Translation (EAMT) . Lisbon, Portugal
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Later among the works it cites.
On the reliability of test collections for evaluating systems of different types. In proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 2101–2104
Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Daniel Campos. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
A comparative study of machine translation for multilingual sentence-level sentiment analysis
Matheus Araújo, Adriano Pereira, and Fabrício Benevenuto. 2020 · 2019
Cited alongside, same era.
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Cited alongside, same era.
Multilingual Transformer Ensembles for Portuguese Natural Language Tasks. In ASSIN@STIL
Ruan Chaves Rodrigues, Jéssica Rodrigues da Silva, Pedro Vitor Quinta de Castro, Nádia Silva, and A. S. Soares. 2019 · 2019
Cited alongside, same era.
PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
Diedre Carmo, Marcos Piau, Israel Campiotti, Rodrigo Nogueira, and Roberto Lotufo. 2020 · 2020
Cited alongside, same era.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H Clark, Jennimaria Palomaki, Vitaly Nikolaev, Eunsol Choi, Dan Garrette, Michael Collins, and Tom Kwiatkowski. 2020 · 2020
Cited alongside, same era.
Understanding BERT Rankers Under Distillation. In Proceedings of the 2020 ACM SIGIR on International Conference on Theory of Information Retrieval (ICTIR 2020) . 149–152
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
RepBERT: Contextualized text embeddings for first-stage retrieval
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020 · 2020
Later among the works it cites.
Overview of the TREC 2020 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2021
Closest in time.
Complement Lexical Retrieval Model with Semantic Residual Embeddings. In Advances in Information Retrieval , Djoerd Hiemstra, Marie-Francine Moens, Josiane Mothe, Raffaele Perego, Martin Potthast, and Fabrizio Sebastiani (Eds.). Springer International Publishing, Cham, 146–160
Luyu Gao, Zhuyun Dai, Tongfei Chen, Zhen Fan, Benjamin Van Durme, and Jamie Callan. 2021 · 2021
Closest in time.
Should we Stop Training More Monolingual Models, and Simply Use Machine Translation Instead?
Tim Isbister, Fredrik Carlsson, and Magnus Sahlgren. 2021 · 2021
Closest in time.
Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. In Proceedings of the 44th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2021)
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021 · 2021
Closest in time.
PROP: Pre-training with Representative Words Prediction for Ad-hoc Retrieval. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining . 283–291
Xinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan, Xiang Ji, and Xueqi Cheng. 2021 · 2021
Closest in time.
The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Closest in time.
RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 5835–5847
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.
A cost-benefit analysis of cross-lingual transfer methods
Guilherme Moraes Rosa, Luiz Henrique Bonifacio, Leandro Rodrigues de Souza, Roberto de Alencar Lotufo, and Rodrigo Nogueira. 2021 · 2021
Closest in time.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Closest in time.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2021 · 2021
Closest in time.