Fetching the paper…
Reading the bibliography…
We study the effect of tokenization on gender bias in machine translation, an aspect that has been largely overlooked in previous works.
A distribution-free k-sample test against ordered alternatives
A. R. Jonckheere. 1954 · 1954
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020a · 2004
Earlier work this paper cites.
Multilingual translation with extensible multilingual pretraining and finetuning
Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020 · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
The social impact of natural language processing
Dirk Hovy and Shannon L. Spruit. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
How much does tokenization affect neural machine translation?
Miguel Domingo, Mercedes García-Martínez, Alexandre Helle, Francisco Casacuberta, and Manuel Herranz. 2018 · 2018
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Equalizing gender bias in neural machine translation with word embeddings techniques
Joel Escudé Font and Marta R. Costa-jussà. 2019 · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Reducing gender bias in neural machine translation as a domain adaptation problem
Danielle Saunders and Bill Byrne. 2020 · 2020
Later among the works it cites.
OPUS-MT — Building open translation services for the World
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Later among the works it cites.
Improving massively multilingual neural machine translation and zero-shot translation
Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020 · 2020
Later among the works it cites.
How to split: the effect of word segmentation on gender bias in speech translation
Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020b · 2020
Cited alongside, same era.
Fine-tuning neural machine translation on gender-balanced datasets
Marta R. Costa-jussà and Adrià de Jorge. 2020 · 2020
Cited alongside, same era.
Gender Bias in Machine Translation
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
Auto-debias: Debiasing masked language models with automated biased prompts
Yue Guo, Yi Yang, and Ahmed Abbasi. 2022 · 2022
Later among the works it cites.
Why don’t people use character-level machine translation?
Jindřich Libovický, Helmut Schmid, and Alexander Fraser. 2022 · 2022
Later among the works it cites.