Fetching the paper…
Reading the bibliography…
Pretrained character-level and byte-level language models have been shown to be competitive with popular subword models across a range of Natural Language Processing (NLP) tasks.
Translation: An Advanced Resource Book
B. Hatim and J. Munday. 2004 · 2004
Earlier work this paper cites.
Hindi-to-Urdu machine translation through transliteration
Nadir Durrani, Hassan Sajjad, Alexander Fraser, and Helmut Schmid. 2010 · 2010
Earlier work this paper cites.
Analyzing the use of character-level translation with sparse and noisy datasets
Jörg Tiedemann and Preslav Nakov. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chung, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Character-based neural machine translation
Marta R. Costa-jussà and José A. R. Fonollosa. 2016 · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander Rush. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2017 · 2017
Earlier work this paper cites.
Byte-based neural machine translation
Marta R. Costa-jussà, Carlos Escolano, and José A. R. Fonollosa. 2017 · 2017
Earlier work this paper cites.
Traducción automática basada en caracteres y redes neuronales
Antonio Manuel Larriba Flor. 2017 · 2017
Earlier work this paper cites.
Fully character-level neural machine translation without explicit segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
How grammatical is character-level neural machine translation? assessing MT quality with contrastive translation pairs
Rico Sennrich. 2017 · 2017
Earlier work this paper cites.
Handling long-term dependencies and rare words in low-resource language modelling
Mittul Singh. 2017 · 2017
Earlier work this paper cites.
Revisiting character-based neural machine translation with capacity and compression
Colin Cherry, George Foster, Ankur Bapna, Orhan Firat, and Wolfgang Macherey. 2018 · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Cited alongside, same era.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. 2018 · 2018
Cited alongside, same era.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. 2023 · 2018
Cited alongside, same era.
Saliency-driven word alignment interpretation for neural machine translation
Shuoyang Ding, Hainan Xu, and Philipp Koehn. 2019 · 2019
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
“will you find these shortcuts?” a protocol for evaluating the faithfulness of input salience methods for text classification
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova. 2022 · 2022
Later among the works it cites.
On the effectiveness of quasi character-level models for machine translation
Salvador Carrión-Ponz and Francisco Casacuberta. 2022 · 2022
Later among the works it cites.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Later among the works it cites.
Subword-delimited downsampling for better character-level translation
Lukas Edman, Antonio Toral, and Gertjan van Noord. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One size does not fit all: Comparing NMT representations of different granularities
Nadir Durrani, Fahim Dalvi, Hassan Sajjad, Yonatan Belinkov, and Preslav Nakov. 2019 · 2019
Cited alongside, same era.
Towards understanding neural machine translation with word importance
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, and Shuming Shi. 2019 · 2019
Cited alongside, same era.
CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Word alignment by fine-tuning embeddings on parallel corpora
Zi-Yi Dou and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Investigating the characteristics of a transformer in a few-shot setup: Does freezing layers in RoBERTa help?
Digvijay Ingle, Rishabh Tripathi, Ayush Kumar, Kevin Patel, and Jithendra Vepa. 2022 · 2022
Later among the works it cites.
Why don’t people use character-level machine translation?
Jindřich Libovický, Helmut Schmid, and Alexander Fraser. 2022 · 2022
Later among the works it cites.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022 · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
NLLB Team, Marta Ruiz Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, et al. 2022 · 2022
Later among the works it cites.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Ricardo Rei, José G. C. De Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Charformer: Fast character transformers via gradient-based subword tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2022 · 2022
Later among the works it cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Accelerating transformer inference for translation via parallel decoding
Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, and Emanuele Rodola. 2023 · 2023
Closest in time.
Inseq: An interpretability toolkit for sequence generation models
Gabriele Sarti, Nils Feldhus, Ludwig Sickert, Oskar van der Wal, Malvina Nissim, and Arianna Bisazza. 2023 · 2023
Closest in time.