Fetching the paper…
Reading the bibliography…
Several neural-based metrics have been recently proposed to evaluate machine translation quality.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Bootstrap methods: Another look at the jackknife
B. Efron. 1979 · 1979
Earlier work this paper cites.
Bayesian Methods for Adaptive Models
David John Cameron Mackay. 1992 · 1992
Earlier work this paper cites.
Aleatory and epistemic uncertainty in probability elicitation with an example from hazardous waste management
Stephen C. Hora. 1996 · 1996
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John C. Platt. 1999 · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G. Dietterich. 2000 · 2000
Earlier work this paper cites.
An introduction to the bootstrap
Roger W. Johnson. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Confidence estimation for machine translation
John Blatz, Erin Fitzgerald, George Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza, Alberto Sanchis, and Nicola Ueffing. 2004 · 2004
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Training a sentence-level machine translation confidence measure
Christopher B. Quirk. 2004 · 2004
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana. 2005 · 2005
Earlier work this paper cites.
Using bayesian model averaging to calibrate forecast ensembles
Adrian E. Raftery, Tilmann Gneiting, Fadoua Balabdaoui, and Michael Polakowski. 2005 · 2005
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Notes on the behavior of mc dropout
Francesco Verdoja and Ville Kyrki. 2020 · 2008
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen. 2009 · 2009
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves. 2011 · 2011
Earlier work this paper cites.
Calibrating predictive model estimates to support personalized medicine
Xiaoqian Jiang, Melanie Osl, Jihoon Kim, and Lucila Ohno-Machado. 2011 · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee Whye Teh. 2011 · 2011
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
Document-level translation quality estimation: exploring discourse and pseudo-references
Carolina Scarton and Lucia Specia. 2014 · 2014
Cited alongside, same era.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015 · 2015
Cited alongside, same era.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Exploring prediction uncertainty in machine translation quality estimation
Daniel Beck, Lucia Specia, and Trevor Cohn. 2016 · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Ensemble learning for multi-source neural machine translation
Ekaterina Garmash and Christof Monz. 2016 · 2016
Improving back-translation with uncertainty-based confidence estimation
Shuo Wang, Yang Liu, Chao Wang, Huanbo Luan, and Maosong Sun. 2019 · 2019
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov. 2020 · 2020
Later among the works it cites.
Practical Guidelines for the Use of MQM in Scientific Research on Translation quality
Aljoscha Burchardt and Arle Lommel. 2014 · 2020
Later among the works it cites.
Posterior network: Uncertainty estimation without ood samples via density-based pseudo-counts
Bertrand Charpentier, Daniel Zügner, and Stephan Günnemann. 2020 · 2020
Later among the works it cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2020
Later among the works it cites.
Unsupervised quality estimation for neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
Bayesian segnet: Model uncertainty in deep convolutional encoder-decoder architectures for scene understanding
Alex Kendall, Vijay Badrinarayanan, and Roberto Cipolla. 2017 · 2017
Cited alongside, same era.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal. 2017 · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017 · 2017
Cited alongside, same era.
Modeling source syntax for neural machine translation
Junhui Li, Deyi Xiong, Zhaopeng Tu, Muhua Zhu, Min Zhang, and Guodong Zhou. 2017 · 2017
Cited alongside, same era.
Robustly representing uncertainty in deep neural networks through sampling
Patrick McClure and Nikolaus Kriegeskorte. 2017 · 2017
Cited alongside, same era.
Marina Fomicheva, Shuo Sun, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, and Lucia Specia. 2020 · 2020
Later among the works it cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
Maximizing overall diversity for improved uncertainty estimates in deep ensembles
Siddhartha Jain, Ge Liu, Jonas Mueller, and David Gifford. 2020 · 2020
Later among the works it cites.
Results of the WMT20 metrics shared task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020 · 2020
Later among the works it cites.
Uncertainty in neural networks: Approximately bayesian ensembling
Tim Pearce, Felix Leibfried, and Alexandra Brintrup. 2020 · 2020
Later among the works it cites.
TransQuest at WMT2020: Sentence-level direct assessment
Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov. 2020 · 2020
Later among the works it cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020a · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Are we estimating or guesstimating translation quality?
Shuo Sun, Francisco Guzmán, and Lucia Specia. 2020 · 2020
Later among the works it cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020a · 2020
Later among the works it cites.
Predicting performance for natural language processing tasks
Mengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Closest in time.
Aleatoric and epistemic uncertainty in machine learning : an introduction to concepts and methods
Eyke Huellermeier and Willem Waegeman. 2021 · 2021
Closest in time.
Uncertainty estimation in autoregressive structured prediction
Andrey Malinin and Mark Gales. 2021 · 2021
Closest in time.
Towards more fine-grained and reliable NLP performance prediction
Zihuiwen Ye, Pengfei Liu, Jinlan Fu, and Graham Neubig. 2021 · 2021
Closest in time.