Fetching the paper…
Reading the bibliography…
The softmax layer in neural machine translation is designed to model the distribution over mutually exclusive tokens.
Calibration of encoder decoder models for neural machine translation
Aviral Kumar and Sunita Sarawagi. 2019 · 1903
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer. 1993 · 1993
Earlier work this paper cites.
A statistical mt tutorial workbook
Kevin Knight. 1999 · 1999
Earlier work this paper cites.
Learning multi-label scene classification
Matthew R. Boutell, Jiebo Luo, Xipeng Shen, and Christopher M. Brown. 2004 · 2004
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Multilabel neural networks with applications to functional genomics and text categorization
Min-Ling Zhang and Zhi-Hua Zhou. 2006 · 2006
Earlier work this paper cites.
Multi-label classification: An overview
Grigorios Tsoumakas and Ioannis Katakis. 2007 · 2007
Earlier work this paper cites.
Parallel implementations of word alignment tool
Qin Gao and Stephan Vogel. 2008 · 2008
Earlier work this paper cites.
Aleatory or epistemic? Does it matter?
Armen Der Kiureghian and Ove Ditlevsen. 2009 · 2009
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn. 2009 · 2009
Earlier work this paper cites.
Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel Cer, Michel Galley, Dan Jurafsky, and Christopher D Manning. 2009 · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
A family of computationally efficient and simple estimators for unnormalized statistical models
Miika Pihlaja, Michael Gutmann, and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
HyTER: Meaning-equivalent semantics for translation evaluation
Markus Dreyer and Daniel Marcu. 2012 · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh. 2012 · 2012
Cited alongside, same era.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Cited alongside, same era.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Length bias in encoder decoder models and a case for global conditioning
Pavel Sountsov and Sunita Sarawagi. 2016 · 2016
Bag-of-words as target for neural machine translation
Shuming Ma, Xu Sun, Yizhong Wang, and Junyang Lin. 2018 · 2018
Later among the works it cites.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang. 2018 · 2018
Later among the works it cites.
Analyzing uncertainty in neural machine translation
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
SGM: Sequence generation model for multi-label classification
Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018 · 2018
Later among the works it cites.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
JAX: Composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018 · 2018
Cited alongside, same era.
Self-normalization properties of language modeling
Jacob Goldberger and Oren Melamud. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne. 2019 · 2019
Later among the works it cites.
A review on multi-label learning algorithms
Min-Ling Zhang and Zhi-Hua Zhou. 2014 · 2019
Later among the works it cites.
Is MAP decoding all you need? the inadequacy of the mode in neural machine translation
Bryan Eikema and Wilker Aziz. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020 · 2020
Later among the works it cites.
Uncertainty determines the adequacy of the mode and the tractability of decoding in sequence-to-sequence models
Felix Stahlberg, Ilia Kulikov, and Shankar Kumar. 2022 · 2022
Closest in time.