Fetching the paper…
Reading the bibliography…
Softmax is the de facto standard in modern neural networks for language processing when it comes to normalizing logits.
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis. 1988 · 1988
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A First Course in Order Statistics (Classics in Applied Mathematics)
Barry C. Arnold, N. Balakrishnan, and H. N. Nagaraja. 2008 · 2008
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Earlier work this paper cites.
Decoding with large-scale neural language models improves translation
Ashish Vaswani, Yinggong Zhao, Victoria Fossum, and David Chiang. 2013 · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Ales Tamchyna. 2014 · 2014
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
On the accuracy of self-normalized log-linear models
Jacob Andreas, Maxim Rabinovich, Michael I Jordan, and Dan Klein. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón Fernandez Astudillo. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Hierarchical neural story generation
Machine translation of restaurant reviews: New corpus for domain adaptation and robustness
Alexandre Berard, Ioan Calapodescu, Marc Dymetman, Claude Roux, Jean-Luc Meunier, and Vassilina Nikoulina. 2019 · 2019
Later among the works it cites.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Later among the works it cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Later among the works it cites.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne. 2019 · 2019
Later among the works it cites.
Attention with sparsity regularization for neural machine translation and summarization
Jiajun Zhang, Yang Zhao, Haoran Li, and Chengqing Zong. 2019 · 2019
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018 · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel. 2018 · 2018
Cited alongside, same era.
Sparse and constrained attention for neural machine translation
Chaitanya Malaviya, Pedro Ferreira, and André F. T. Martins. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Is sparse attention more interpretable?
Clara Meister, Stefan Lazov, Isabelle Augenstein, and Ryan Cotterell. 2021 · 2021
Closest in time.
Smoothing and shrinking the sparse seq2seq search space
Ben Peters and André F. T. Martins. 2021 · 2021
Closest in time.
Sparse attention with linear units
Biao Zhang, Ivan Titov, and Rico Sennrich. 2021 · 2021
Closest in time.