Fetching the paper…
Reading the bibliography…
The choice of token vocabulary affects the performance of machine translation.
A note on measurement of utility
Paul A Samuelson. 1937 · 1937
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Another look at the data sparsity problem
Ben Allison, David Guthrie, and Louise Guthrie. 2006 · 2006
Earlier work this paper cites.
Very deep transformers for neural machine translation
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao. 2020 · 2008
Earlier work this paper cites.
Mathematical theory of entropy
Nathaniel FG Martin and James W England. 2011 · 2011
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
The word entropy of natural languages
Christian Bentz and Dimitrios Alikaniotis. 2016 · 2016
Earlier work this paper cites.
Character-based neural machine translation
Marta R. Costa-jussà and José A. R. Fonollosa. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Fully character-level neural machine translation without explicit segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Revisiting character-based neural machine translation with capacity and compression
Colin Cherry, George F. Foster, Ankur Bapna, Orhan Firat, and Wolfgang Macherey. 2018 · 2018
Cited alongside, same era.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush. 2018 · 2018
Cited alongside, same era.
Learning to segment inputs for NMT favors character-level processing
Julia Kreutzer and Artem Sokolov. 2018 · 2018
Cited alongside, same era.
How large a vocabulary does text classification need? A variational approach to vocabulary selection
Wenhu Chen, Yu Su, Yilin Shen, Zhiyu Chen, Xifeng Yan, and William Yang Wang. 2019 · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
A call for prudent choice of subword merge operations in neural machine translation
Shuoyang Ding, Adithya Renduchintala, and Kevin Duh. 2019 · 2019
Later among the works it cites.
Computational optimal transport
Gabriel Peyré and Marco Cuturi. 2019 · 2019
Later among the works it cites.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Later among the works it cites.
The evolved transformer
David R. So, Quoc V. Le, and Chen Liang. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Singh Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Deep learning for sentiment analysis: A survey
Lei Zhang, Shuai Wang, and Bing Liu. 2018 · 2018
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Recurrent neural network for text classification with hierarchical multiscale dense connections
Yi Zhao, Yanyan Shen, and Junjie Yao. 2019 · 2019
Later among the works it cites.
Bpe-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Closest in time.
Optimizing segmentation granularity for neural machine translation
Elizabeth Salesky, Andrew Runge, Alex Coda, Jan Niehues, and Graham Neubig. 2020 · 2020
Closest in time.
Neural machine translation with byte-level subwords
Changhan Wang, Kyunghyun Cho, and Jiatao Gu. 2020 · 2020
Closest in time.