Fetching the paper…
Reading the bibliography…
Gender-inclusive NLP research has documented the harmful limitations of gender binary-centric large language models (LLM), such as the inability to correctly use gender-diverse English neopronouns (e.g., xe, zir, fae).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021 · 1994
Earlier work this paper cites.
Historical Linguistics 2003: Selected Papers from the 16th International Conference on Historical Linguistics, Copenhagen, 11-15 August 2003
M.D. Fortescue. 2005 · 2003
Earlier work this paper cites.
Morphology
Penny Eckert and Ivan A. Sag. 2011 · 2011
Earlier work this paper cites.
The Chicago Guide to Grammar, Usage, and Punctuation
B.A. Garner. 2016 · 2016
Earlier work this paper cites.
Nounself pronouns: 3rd person personal pronouns as identity expression
Ehm Hjorth Miltersen. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
How much does tokenization affect neural machine translation?
Miguel Domingo, Mercedes García-Martínez, Alexandre Helle, Francisco Casacuberta, and Manuel Herranz. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Better oov translation with bilingual terminology mining
Matthias Huck, Viktor Hangya, and Alexander M. Fraser. 2019 · 2019
Earlier work this paper cites.
Man is to person as woman is to location: Measuring gender bias in named entity recognition
Ninareh Mehrabi, Thamme Gowda, Fred Morstatter, Nanyun Peng, and A. G. Galstyan. 2019 · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A Smith, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Mitigating gender bias in natural language processing: Literature review
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Improving pre-trained multilingual model with vocabulary expansion
Hai Wang, Dian Yu, Kai Sun, Jianshu Chen, and Dong Yu. 2019 · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Quality of word vectors and its impact on named entity recognition in czech
František Dařena and Martin Süss. 2020 · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Measuring harmful sentence completion in language models for lgbtqia+ individuals
Debora Nozza, Federico Bianchi, Anne Lauscher, Dirk Hovy, et al. 2022 · 2022
Later among the works it cites.
Back to the future: On potential histories in nlp
Zeerak Talat and Anne Lauscher. 2022 · 2022
Later among the works it cites.
Miner: Improving out-of-vocabulary named entity recognition from an information theoretic perspective
Xiao Wang, Shihan Dou, Li Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, and Xuanjing Huang. 2022 · 2022
Later among the works it cites.
Incorporating context into subword vocabularies
Shaked Yehezkel and Yuval Pinter. 2022 · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Rose Biderman, Hailey Schoelkopf, Quentin G. Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
As good as new. how to successfully recycle english gpt-2 to make models for other languages
Wietse de Vries and Malvina Nissim. 2021 · 2021
Cited alongside, same era.
How to split: the effect of word segmentation on gender bias in speech translation
Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Cited alongside, same era.
Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2021 · 2021
Cited alongside, same era.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. 2021 · 2021
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
A survey on gender bias in natural language processing
Karolina Stanczak and Isabelle Augenstein. 2021 · 2021
Cited alongside, same era.
Closest in time.
Winoqueer: A community-in-the-loop benchmark for anti-lgbtq+ bias in large language models
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023 · 2023
Closest in time.
2023 gender census
Gender Census. 2023 · 2023
Closest in time.
Misgendered: Limits of large language models in understanding pronouns
Tamanna Hossain, Sunipa Dev, and Sameer Singh. 2023 · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Closest in time.
Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages
Tomasz Limisiewicz, Jiří Balhar, and David Mareček. 2023 · 2023
Closest in time.
Bound by the bounty: Collaboratively shaping evaluation processes for queer ai harms
Organizers of QueerInAI, Nathaniel Dennler, Anaelia Ovalle, Ashwin Singh, Luca Soldaini, Arjun Subramonian, Huy Tu, William Agnew, Avijit Ghosh, Kyra Yee, Irene Font Peradejordi, Zeerak Talat, Mayra Russo, and Jessica de Jesus de Pinho Pinhal. 2023 · 2023
Closest in time.
“i’m fully who i am”: Towards centering transgender and non-binary voices to measure biases in open language generation
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.