Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are trained on corpora disproportionally weighted in favor of Standard American English.
Earth mover’s distance minimization for unsupervised bilingual lexicon induction
Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 1945
Earlier work this paper cites.
Optimal transport: Old and new
Cédric Villani. 2008 · 2008
Earlier work this paper cites.
Data-driven dialectology
John Nerbonne. 2009 · 2009
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi. 2013 · 2013
Earlier work this paper cites.
Automatically processing tweets from gang-involved youth: Towards detecting loss and aggression
Terra Blevins, Robert Kwiatkowski, Jamie MacBeth, Kathleen McKeown, Desmond Patton, and Owen Rambow. 2016 · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V. Le. 2016 · 2016
Earlier work this paper cites.
The social impact of natural language processing
Dirk Hovy and Shannon L. Spruit. 2016 · 2016
Earlier work this paper cites.
Learning a POS tagger for AAVE-like language
Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2016 · 2016
Earlier work this paper cites.
Martin Arjovsky, Soumith Chintala, and Léon Bottou. 2017 · 2017
Earlier work this paper cites.
Su Lin Blodgett and Brendan O’Connor. 2017 · 2017
Earlier work this paper cites.
Incorporating dialectal variability for socially equitable language identification
David Jurgens, Yulia Tsvetkov, and Dan Jurafsky. 2017 · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2017 · 2017
Earlier work this paper cites.
Gromov-Wasserstein alignment of word embedding spaces
David Alvarez-Melis and Tommi Jaakkola. 2018 · 2018
Earlier work this paper cites.
Twitter Universal Dependency parsing for African-American and mainstream American English
Su Lin Blodgett, Johnny Wei, and Brendan O’Connor. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Interpolating between optimal transport and mmd using sinkhorn divergences
Jean Feydy, Thibault Séjourné, François-Xavier Vialard, Shun ichi Amari, Alain Trouvé, and Gabriel Peyré. 2018 · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif Mohammad. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Parameter prediction for unseen deep architectures
Boris Knyazev, Michal Drozdzal, Graham W. Taylor, and Adriana Romero-Soriano. 2021 · 2021
Later among the works it cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021 · 2021
Later among the works it cites.
Challenges in automated debiasing for toxic language detection
Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019 · 2019
Cited alongside, same era.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Adversarial decomposition of text representation
Alexey Romanov, Anna Rumshisky, Anna Rogers, and David Donahue. 2019 · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Steven Bird. 2022 · 2022
Later among the works it cites.
Whose language counts as high quality? measuring language ideologies in text data selection
Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A. Smith. 2022 · 2022
Later among the works it cites.
Hyperprompt: Prompt-based task-conditioning of transformers
Yun He, Huaixiu Steven Zheng, Yi Tay, Jai Gupta, Yu Du, Vamsi Aribandi, Zhe Zhao, YaGuang Li, Zhao Chen, Donald Metzler, Heng-Tze Cheng, and Ed H. Chi. 2022 · 2022
Later among the works it cites.
Hypertuning: Toward adapting large language models without back-propagation
Jason Phang, Yi Mao, Pengcheng He, and Weizhu Chen. 2022 · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2022 · 2022
Later among the works it cites.
VALUE: Understanding dialect disparity in NLU
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Hyper-x: A unified hypernetwork for multi-task multilingual transfer
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord, and Sebastian Ruder. 2022 · 2022
Later among the works it cites.
Tada: Task-agnostic dialect adapters for english
Will Held, Caleb Ziems, and Diyi Yang. 2023 · 2023
Closest in time.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. 2023 · 2023
Closest in time.
Dada: Dialect adaptation via dynamic aggregation of linguistic rules
Yanchen Liu, William Held, and Diyi Yang. 2023 · 2023
Closest in time.
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić, and Edoardo Maria Ponti. 2023 · 2023
Closest in time.