Fetching the paper…
Reading the bibliography…
Language models perform differently across languages.
Cognitive prerequisites for the development of grammar
Dan I Slobin. 1973 · 1973
Earlier work this paper cites.
The Icelandic parsed historical corpus (IcePaHC)
Eiríkur Rögnvaldsson, Anton Karl Ingason, Einar Freyr Sigur \textipa · 1984
Earlier work this paper cites.
Of abundance and scantiness in inflection: A typological prelude
Frans Plank. 1991 · 1991
Earlier work this paper cites.
The morphology of the mental lexicon: Internal word structure viewed from a psycholinguistic perspective
Dominiek Sandra. 1994 · 1994
Earlier work this paper cites.
The polysynthesis parameter
Mark Baker. 1996 · 1996
Earlier work this paper cites.
Cognition, quantitative linguistics, and systemic typology
Gertraud Fenk-Oczlon and August Fenk. 1999 · 1999
Earlier work this paper cites.
Split morphology: How agglutination and flexion mix
Frans Plank. 1999 · 1999
Earlier work this paper cites.
Form-function relations: how do children find out what they are?
Dan I Slobin. 2001 · 2001
Earlier work this paper cites.
Statistical morphological disambiguation for agglutinative languages
Dilek Z Hakkani-Tür, Kemal Oflazer, and Gökhan Tür. 2002 · 2002
Earlier work this paper cites.
Design and implementation of the bulgarian hpsg-based treebank
Kiril Simov, Petya Osenova, Alexander Simov, and Milen Kouylekov. 2005 · 2005
Earlier work this paper cites.
AnCora: Multilevel annotated corpora for Catalan and Spanish
Mariona Taulé, M. Antònia Martí, and Marta Recasens. 2008 · 2008
Earlier work this paper cites.
An empirical test of the Agglutination Hypothesis
Martin Haspelmath. 2009 · 2009
Earlier work this paper cites.
Hindi syntax: Annotating dependency, lexical predicate-argument structure, and phrase structure
Martha Palmer, Rajesh Bhatt, Bhuvana Narasimhan, Owen Rambow, Dipti Misra Sharma, and Fei Xia. 2009 · 2009
Earlier work this paper cites.
Hybrid n-gram probability estimation in morphologically rich languages
Hyopil Shin and Hyunjo You. 2009 · 2009
Earlier work this paper cites.
lme4: Mixed-effects modeling with R
Douglas M Bates. 2010 · 2010
Earlier work this paper cites.
Morphological typology
Dunstan Patrick Brown. 2010 · 2010
Earlier work this paper cites.
A typological approach to first language acquisition
Wolfgang U Dressler. 2010 · 2010
Earlier work this paper cites.
Hungarian dependency treebank
Veronika Vincze, Dóra Szauter, Attila Almási, György Móra, Zoltán Alexin, and János Csirik. 2010 · 2010
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Emily M Bender. 2011 · 2011
Earlier work this paper cites.
Indonesian morphology tool (morphind): Towards an indonesian corpus
Septina Dian Larasati, Vladislav Kuboň, and Daniel Zeman. 2011 · 2011
Earlier work this paper cites.
Prague dependency style treebank for Tamil
Loganathan Ramasamy and Zdeněk Žabokrtský. 2012 · 2012
Earlier work this paper cites.
Morphological organization: The low conditional entropy conjecture
Farrell Ackerman and Robert Malouf. 2013 · 2013
Earlier work this paper cites.
WALS Online (v2020.3)
Matthew S. Dryer and Martin Haspelmath. 2013 · 2013
Earlier work this paper cites.
Development of a Persian syntactic dependency treebank
Mohammad Sadegh Rasooli, Manouchehr Kouhestani, and Amirsaeid Moloodi. 2013 · 2013
Earlier work this paper cites.
Crosslinguistic evidence for the language-making capacity
Dan I Slobin. 2013 · 2013
Earlier work this paper cites.
Experiments for dependency parsing of Greek
Prokopis Prokopidis and Haris Papageorgiou. 2014 · 2014
Earlier work this paper cites.
Automatic conversion of the basque dependency treebank to universal dependencies
Maria Jesus Aranzabe, Aitziber Atutxa, Kepa Bengoetxea, Koldo Gojenola, and Larraitz Uria. 2015 · 2015
Earlier work this paper cites.
Universal Dependencies for Irish
Teresa Lynn and Jennifer Foster. 2016 · 2016
Earlier work this paper cites.
The hindi/urdu treebank project
Riyaz Ahmad Bhat, Rajesh Bhatt, Annahita Farudi, Prescott Klassen, Bhuvana Narasimhan, Martha Palmer, Owen Rambow, Dipti Misra Sharma, Ashwini Vaidya, Sri Ramagurumurthy Vishnu, et al. 2017 · 2017
Earlier work this paper cites.
The Universal Dependencies treebank for Slovenian
Kaja Dobrovoljc, Tomaž Erjavec, and Simon Krek. 2017 · 2017
Earlier work this paper cites.
Neural machine translation for morphologically rich languages with improved sub-word units and synthetic data
Mārcis Pinnis, Rihards Krišlauks, Daiga Deksne, and Toms Miks. 2017 · 2017
Earlier work this paper cites.
The GUM corpus: Creating multilayer resources in the classroom
Amir Zeldes. 2017 · 2017
Earlier work this paper cites.
Meaningless yet meaningful: Morphology grounded subword-level NMT
Tamali Banerjee and Pushpak Bhattacharyya. 2018 · 2018
Earlier work this paper cites.
Building Universal Dependency treebanks in Korean
Jayeol Chun, Na-Rae Han, Jena D. Hwang, and Jinho D. Choi. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Are all languages equally hard to language-model?
Ryan Cotterell, Sabrina J. Mielke, Jason Eisner, and Brian Roark. 2018 · 2018
Cited alongside, same era.
On the relation between linguistic typology and (limitations of) multilingual language modeling
Daniela Gerz, Ivan Vulić, Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen. 2018b · 2018
Cited alongside, same era.
UniMorph 2.0: Universal Morphology
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia, Manaal Faruqui, Sabrina J. Mielke, Arya McCarthy, Sandra Kübler, David Yarowsky, Jason Eisner, and Mans Hulden. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Morphological and language-agnostic word segmentation for NMT
Dominik Macháček, Jonáš Vidra, and Ondřej Bojar. 2018 · 2018
Cited alongside, same era.
When Is Multilinguality a Curse? Language Modeling for 250 High-and Low-Resource Languages
Tyler A Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K Bergen. 2023 · 2023
Later among the works it cites.
A study on the evaluation of tokenizer performance in natural language processing
Sanghyun Choo and Wonjoon Kim. 2023 · 2023
Later among the works it cites.
Generative AI has a language problem
Monojit Choudhury. 2023 · 2023
Later among the works it cites.
Explicit Morphological Knowledge Improves Pre-training of Language Models for Hebrew
Eylon Gueta, Omer Goldman, and Reut Tsarfaty. 2023 · 2023
Later among the works it cites.
Languages through the looking glass of BPE compression
Ximena Gutierrez-Vasques, Christian Bentz, and Tanja Samardžić. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Morphological zero-shot neural machine translation
Giulio Zhou. 2018 · 2018
Cited alongside, same era.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. 2019 · 2019
Cited alongside, same era.
Morphology-aware word-segmentation in dialectal Arabic adaptation of neural machine translation
Ahmed Tawfik, Mahitab Emam, Khaled Essam, Robert Nabil, and Hany Hassan. 2019 · 2019
Cited alongside, same era.
Building language models for morphological rich low-resource languages using data from related donor languages: the case of Uyghur
Ayimunishagu Abulimiti and Tanja Schultz. 2020 · 2020
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Wiki-40B: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Haris Jabbar. 2023 · 2023
Later among the works it cites.
Evaluating the diversity, equity, and inclusion of NLP technology: A case study for Indian languages
Simran Khanuja, Sebastian Ruder, and Partha Talukdar. 2023 · 2023
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023 · 2023
Later among the works it cites.
XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
CompoundPiece: Evaluating and improving decompounding performance of language models
Benjamin Minixhofer, Jonas Pfeiffer, and Ivan Vulić. 2023 · 2023
Later among the works it cites.
Crosslingual Generalization through Multitask Finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2023 · 2023
Later among the works it cites.
Fairness in language models beyond English: Gaps and challenges
Krithika Ramesh, Sunayana Sitaram, and Monojit Choudhury. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
SIB-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects
David Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba Alabi, Yanke Mao, Haonan Gao, and En-Shiun Lee. 2024 · 2024
Closest in time.
Ud cebuano-gja
Aranes, Glyd Jun and Zeman, Dan . 2021 · 2024
Closest in time.
A bit of a problem: Measurement disparities in dataset sizes across languages
Catherine Arnett, Tyler A. Chang, and Benjamin Bergen. 2024a · 2024
Closest in time.
BPE-knockout: Pruning pre-existing BPE tokenisers with backwards-compatible morphological semi-supervision
Thomas Bauwens and Pieter Delobelle. 2024 · 2024
Closest in time.
Goldfish: Monolingual Language Models for 350 Languages
Tyler A Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K Bergen. 2024 · 2024
Closest in time.
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, and Ivan P. Yamshchikov. 2024 · 2024
Closest in time.
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Roziere. 2024 · 2024
Closest in time.
Language modeling is compression
Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness. 2024 · 2024
Closest in time.
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024 · 2024
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent Sifre. 2024 · 2024
Closest in time.
Mission: Impossible language models
Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, and Christopher Potts. 2024 · 2024
Closest in time.
Effect of tokenization granularity for Turkish large language models
Yiğit Bekir Kaya and A Cüneyd Tantuğ. 2024 · 2024
Closest in time.
Length-aware byte pair encoding for mitigating over-segmentation in Korean machine translation
Jungseob Lee, Hyeonseok Moon, Seungjun Lee, Chanjun Park, Sugyeong Eo, Hyunwoong Ko, Jaehyung Seo, Seungyoon Lee, and Heuiseok Lim. 2024 · 2024
Closest in time.
Lexically grounded subword segmentation
Jindřich Libovickỳ and Jindřich Helcl. 2024 · 2024
Closest in time.
MaLA-500: Massive Language Adaptation of Large Language Models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André FT Martins, and Hinrich Schütze. 2024 · 2024
Closest in time.
J-UniMorph: Japanese morphological annotation through the universal feature schema
Kosuke Matsuzaki, Masaya Taniguchi, Kentaro Inui, and Keisuke Sakaguchi. 2024 · 2024
Closest in time.
Tokenization is more than compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024 · 2024
Closest in time.
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
Omri Uzan, Craig W Schmidt, Chris Tanner, and Yuval Pinter. 2024 · 2024
Closest in time.
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Generative LLM Inference
Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras. 2024 · 2024
Closest in time.
Fast WordPiece Tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, and Denny Zhou. 2021 · 2089
Closest in time.