Fetching the paper…
Reading the bibliography…
A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Multi-stage document ranking with bert
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019 · 1910
Earlier work this paper cites.
WIKIR: A python toolkit for building a large-scale Wikipedia-based English information retrieval dataset
Jibril Frej, Didier Schwab, and Jean-Pierre Chevallet. 2020 · 1933
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Die Einteilung der deutschen Dialekte
Peter Wiesinger. 1983 · 1983
Earlier work this paper cites.
Variation im Deutschen: Soziolinguistische Perspektiven
Stephen Barbour and Patrick Stevenson. 1998 · 1998
Earlier work this paper cites.
Bridging the lexical chasm: statistical approaches to answer-finding
Adam Berger, Rich Caruana, David Cohn, Dayne Freitag, and Vibhu Mittal. 2000 · 2000
Earlier work this paper cites.
Die nordfriesischen Mundarten
Alastair G. H. Walker and Ommo Wilts. 2001 · 2001
Earlier work this paper cites.
A history of twentieth-century american academic cartography
Robert McMaster and Susanna McMaster. 2002 · 2002
Earlier work this paper cites.
Clef 2003 – overview of results
Martin Braschler. 2004 · 2003
Earlier work this paper cites.
Atlas zur deutschen Alltagssprache (AdA)
Stephan Elspaß and Robert Möller. 2003 · 2003
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al. 2009 · 2009
Earlier work this paper cites.
The tower of babel meets web 2.0: user-generated content and its applications in a multilingual context
Brent Hecht and Darren Gergle. 2010 · 2010
Earlier work this paper cites.
Learning translational and knowledge-based similarities from relevance rankings for cross-language retrieval
Shigehiko Schamoni, Felix Hieber, Artem Sokolov, and Stefan Riezler. 2014 · 2014
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016 · 2016
Earlier work this paper cites.
Detecting cross-cultural differences using a multilingual topic model
E.D. Gutiérrez, Ekaterina Shutova, Patricia Lichtenstein, Gerard de Melo, and Luca Gilardi. 2016 · 2016
Earlier work this paper cites.
Cross-lingual learning-to-rank with shared representations
Shota Sasaki, Shuo Sun, Shigehiko Schamoni, Kevin Duh, and Kentaro Inui. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The challenges of optimizing machine translation for low resource cross-language information retrieval
Constantine Lignos, Daniel Cohen, Yen-Chieh Lien, Pratik Mehta, W. Bruce Croft, and Scott Miller. 2019 · 2019
Earlier work this paper cites.
Cedr: Contextualized embeddings for document ranking
Sean MacAvaney, Andrew Yates, Arman Cohan, and Nazli Goharian. 2019 · 2019
Earlier work this paper cites.
Unsupervised data augmentation for less-resourced languages with no standardized spelling
Alice Millour and Karën Fort. 2019 · 2019
Cited alongside, same era.
The effect of translationese in machine translation test sets
Mike Zhang and Antonio Toral. 2019 · 2019
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Omar Khattab and Matei Zaharia. 2020 · 2020
Cited alongside, same era.
On cross-lingual retrieval with multilingual text encoders
Robert Litschko, Ivan Vulić, Simone Paolo Ponzetto, and Goran Glavaš. 2022 · 2022
Later among the works it cites.
Streamlining evaluation with ir-measures
Sean MacAvaney, Craig Macdonald, and Iadh Ounis. 2022 · 2022
Later among the works it cites.
AfriCLIRMatrix: Enabling cross-lingual information retrieval for African languages
Odunayo Ogundepo, Xinyu Zhang, Shuo Sun, Kevin Duh, and Jimmy Lin. 2022 · 2022
Later among the works it cites.
ColBERTv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Later among the works it cites.
DEPTH+: An enhanced depth metric for Wikipedia corpora quality
Saied Alshahrani, Norah Alshahrani, and Jeanna Matthews. 2023 · 2023
Later among the works it cites.
Low-resource bilingual dialect lexicon induction with large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Teaching a new dog old tricks: Resurrecting multilingual retrieval using zero-shot learning
Sean MacAvaney, Luca Soldaini, and Nazli Goharian. 2020 · 2020
Cited alongside, same era.
LAReQA: Language-agnostic answer retrieval from a multilingual pool
Uma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua, Aaron Phillips, and Yinfei Yang. 2020 · 2020
Cited alongside, same era.
Document translation vs. query translation for cross-lingual information retrieval in the medical domain
Shadi Saleh and Pavel Pecina. 2020 · 2020
Cited alongside, same era.
Cross-lingual training of neural models for document ranking
Peng Shi, He Bai, and Jimmy Lin. 2020 · 2020
Cited alongside, same era.
CLIRMatrix: A massively large collection of bilingual and multilingual datasets for cross-lingual information retrieval
Shuo Sun and Kevin Duh. 2020 · 2020
Cited alongside, same era.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Cited alongside, same era.
Ekaterina Artemova and Barbara Plank. 2023 · 2023
Later among the works it cites.
Revisiting machine translation for cross-lingual classification
Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
On the effects of regional spelling conventions in retrieval models
Andreas Chari, Sean MacAvaney, and Iadh Ounis. 2023 · 2023
Later among the works it cites.
Boosting zero-shot cross-lingual retrieval by training on artificially code-switched data
Robert Litschko, Ekaterina Artemova, and Barbara Plank. 2023 · 2023
Later among the works it cites.
Is ChatGPT good at search? investigating large language models as re-ranking agents
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Later among the works it cites.
Zero-shot cross-lingual reranking with large language models for low-resource languages
Mofetoluwa Adeyemi, Akintunde Oladipo, Ronak Pradeep, and Jimmy Lin. 2024 · 2024
Closest in time.
Orphan articles: The dark matter of wikipedia
Akhil Arora, Robert West, and Martin Gerlach. 2024 · 2024
Closest in time.
DIALECTBENCH: An NLP benchmark for dialects, varieties, and closely-related languages
Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, and Antonios Anastasopoulos. 2024 · 2024
Closest in time.
Wikipedia article depth
Meta contributors. 2024 · 2024
Closest in time.
Kardeş-NLU: Transfer to low-resource languages with the help of a high-resource cousin – a benchmark and evaluation for Turkic languages
Lütfi Kerem Senel, Benedikt Ebing, Konul Baghirova, Hinrich Schuetze, and Goran Glavaš. 2024 · 2024
Closest in time.
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, and Miryam de Lhoneux. 2024 · 2024
Closest in time.
Messirve: A large-scale spanish information retrieval dataset
Francisco Valentini, Viviana Cotik, Damián Furman, Ivan Bercovich, Edgar Altszyler, and Juan Manuel Pérez. 2024 · 2024
Closest in time.
Modular adaptation of multilingual encoders to written Swiss German dialect
Jannis Vamvas, Noëmi Aepli, and Rico Sennrich. 2024 · 2024
Closest in time.
Creating a lexicon of Bavarian dialect by means of Facebook language data and crowdsourcing
Manuel Burghardt, Daniel Granvogl, and Christian Wolff. 2016 · 2033
Closest in time.