Fetching the paper…
Reading the bibliography…
Current sequence-to-sequence models are trained to minimize cross-entropy and use softmax to compute the locally normalized probabilities over target sequences.
Calibration of encoder decoder models for neural machine translation
Aviral Kumar and Sunita Sarawagi. 2019 · 1903
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler. 1951 · 1951
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort
D Raj Reddy et al. 1977 · 1977
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis. 1988 · 1988
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
John S Bridle. 1990 · 1990
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Clustering with bregman divergences
Arindam Banerjee, Srujana Merugu, Inderjit S Dhillon, and Joydeep Ghosh. 2005 · 2005
Earlier work this paper cites.
Is map decoding all you need? the inadequacy of the mode in neural machine translation
Bryan Eikema and Wilker Aziz. 2020 · 2005
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J Wainwright and Michael I Jordan. 2008 · 2008
Earlier work this paper cites.
Convex analysis and nonlinear optimization: theory and examples
Jonathan Borwein and Adrian S Lewis. 2010 · 2010
Earlier work this paper cites.
The Kyoto free translation task
Graham Neubig. 2011 · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Sequence-to-sequence neural net models for grapheme-to-phoneme conversion
Kaisheng Yao and Geoffrey Zweig. 2015 · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Morphological inflection generation using character sequence to sequence learning
Manaal Faruqui, Yulia Tsvetkov, Graham Neubig, and Chris Dyer. 2016 · 2016
Cited alongside, same era.
Improved neural machine translation with smt features
Wei He, Zhongjun He, Hua Wu, and Haifeng Wang. 2016 · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André FT Martins and Ramón Fernandez Astudillo. 2016 · 2016
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Later among the works it cites.
Joey NMT: A minimalist NMT toolkit for novices
Julia Kreutzer, Jasmijn Bastings, and Stefan Riezler. 2019 · 2019
Later among the works it cites.
The SIGMORPHON 2019 shared task: Morphological analysis in context and cross-lingual transfer for inflection
Arya D. McCarthy, Ekaterina Vylomova, Shijie Wu, Chaitanya Malaviya, Lawrence Wolf-Sonkin, Garrett Nicolai, Christo Kirov, Miikka Silfverberg, Sabrina J. Mielke, Jeffrey Heinz, Ryan Cotterell, and Mans Hulden. 2019 · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. 2019 · 2019
Later among the works it cites.
IT–IST at the SIGMORPHON 2019 shared task: Sparse two-headed models for inflection
Ben Peters and André F. T. Martins. 2019 · 2019
Later among the works it cites.
Sparse sequence-to-sequence models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Overview of the IWSLT 2017 evaluation campaign
M Cettolo, M Federico, L Bentivogli, J Niehues, S Stüker, K Sudoh, K Yoshino, and C Federmann. 2017 · 2017
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton. 2017 · 2017
Cited alongside, same era.
Ben Peters, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Later among the works it cites.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne. 2019 · 2019
Later among the works it cites.
Learning with fenchel-young losses
Mathieu Blondel, André FT Martins, and Vlad Niculae. 2020 · 2020
Later among the works it cites.
SIGMORPHON 2020 task 0 system description: ETH Zürich team
Martina Forster and Clara Meister. 2020 · 2020
Later among the works it cites.
The sigmorphon 2020 shared task on multilingual grapheme-to-phoneme conversion
Kyle Gorman, Lucas F.E. Ashby, Aaron Goyzueta, Arya D. McCarthy, Shijie Wu, and Daniel You. 2020 · 2020
Later among the works it cites.
Semantic label smoothing for sequence to sequence problems
Michal Lukasik, Himanshu Jain, Aditya Menon, Seungyeon Kim, Srinadh Bhojanapalli, Felix Yu, and Sanjiv Kumar. 2020 · 2020
Later among the works it cites.
If beam search is the answer, what was the question?
Clara Meister, Ryan Cotterell, and Tim Vieira. 2020a · 2020
Later among the works it cites.
One-size-fits-all multilingual models
Ben Peters and André F. T. Martins. 2020 · 2020
Later among the works it cites.
On the inference calibration of neural machine translation
Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020 · 2020
Later among the works it cites.
Xiang Yu, Ngoc Thang Vu, and Jonas Kuhn. 2020 · 2020
Later among the works it cites.
Token-level and sequence-level loss smoothing for RNN language models
Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2018 · 2094
Closest in time.