Fetching the paper…
Reading the bibliography…
This paper presents a new approach for assessing uncertainty in machine translation by simultaneously evaluating translation quality and providing a reliable confidence score.
Learning by transduction
A. Gammerman, V. Vovk, and V. Vapnik · 1998
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Algorithmic learning in a random world
Vladimir Vovk, Alex Gammerman, and Glenn Shafer · 2005
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul · 2006
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael J Denkowski · 2009
Earlier work this paper cites.
Findings of the 2012 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia · 2012
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel · 2013
Earlier work this paper cites.
Venn-abers predictors
V. Vovk and Ivan Petej · 2014
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Exploring prediction uncertainty in machine translation quality estimation
Daniel Beck, Lucia Specia, and Trevor Cohn · 2016
Earlier work this paper cites.
Nonparametric predictive distributions based on conformal prediction
Vladimir Vovk, Jieli Shen, Valery Manokhin, and Min-ge Xie · 2017
Earlier work this paper cites.
Accurate uncertainties for deep learning using calibrated regression
Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
A deep neural network conformal predictor for multi-label text classification
Andreas Paisios, Ladislav Lenc, Jiří Martínek, Pavel Král, and Harris Papadopoulos · 2019
Cited alongside, same era.
Computationally efficient versions of conformal predictive distributions
Vladimir Vovk, Ivan Petej, Ilia Nouretdinov, Valery Manokhin, and Alexander Gammerman · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Transformer-based conformal predictors for paraphrase detection
Patrizio Giovannotti and Alex Gammerman · 2021
Later among the works it cites.
Uncertainty-aware machine translation evaluation
Taisiya Glushkova, Chrysoula Zerva, Ricardo Rei, and André F. T. Martins · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf · 2021
Later among the works it cites.
Retrain or not retrain: conformal test martingales for change-point detection
Vladimir Vovk, Ivan Petej, Ilia Nouretdinov, Ernst Ahlberg, Lars Carlsson, and Alex Gammerman · 2021
Later among the works it cites.
A global analysis of metrics used for measuring performance in natural language processing
Kathrin Blagec, Georg Dorffner, Milad Moradi, Simon Ott, and Matthias Samwald · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert-based conformal predictor for sentiment analysis
Lysimachos Maltoudoglou, Andreas Paisios, and Harris Papadopoulos · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Cited alongside, same era.
Findings of the WMT 2020 shared task on quality estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Erick Fonseca, Vishrav Chaudhary, Francisco Guzmán, and André F. T. Martins · 2020
Cited alongside, same era.
Tutorial on conformal predictive distributions, 2020
Paolo Toccaceli · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Later among the works it cites.
crepes: a python package for generating conformal regressors and predictive systems
Henrik Boström · 2022
Later among the works it cites.
Conformal prediction for text infilling and part-of-speech prediction
Neil Dey, Jing Ding, Jack Ferrell, Carolina Kapper, Maxwell Lovig, Emiliano Planchon, and Jonathan P. Williams · 2022
Later among the works it cites.
Calibration of natural language understanding models with venn–abers predictors
Patrizio Giovannotti · 2022
Later among the works it cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler · 2022
Later among the works it cites.
Layer or representation space: What makes BERT-based evaluation metrics robust?
Doan Nam Long Vu, Nafise Sadat Moosavi, and Steffen Eger · 2022
Later among the works it cites.
Conformal nucleus sampling, 2023
Shauli Ravfogel, Yoav Goldberg, and Jacob Goldberger · 2023
Closest in time.
Leveraging large language models for multiple choice question answering, 2023
Joshua Robinson, Christopher Michael Rytting, and David Wingate · 2023
Closest in time.