Fetching the paper…
Reading the bibliography…
Traditional automatic evaluation measures for natural language generation (NLG) use costly human-authored references to estimate the quality of a system output.
Regression analysis
Williams, Evan James · 1959
Earlier work this paper cites.
The measurement of observer agreement for categorical data
Landis, J Richard and Koch, Gary G · 1977
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
A Neural Probabilistic Language Model
Bengio, Yoshua, Ducharme, Réjean, Vincent, Pascal, and Jauvin, Christian · 2003
Earlier work this paper cites.
How does automatic machine translation evaluation correlate with human scoring as the number of reference translations increases?
Finch, Andrew M, Akiba, Yasuhiro, and Sumita, Eiichiro · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, Chin-Yew · 2004
Earlier work this paper cites.
Meteor: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments
Lavie, Alon and Agarwal, Abhaya · 2007
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
Mairesse, F., Gašić, M., Jurčíček, F., Keizer, S., Thomson, B., Yu, K., and Young, S · 2010
Earlier work this paper cites.
Machine translation evaluation versus quality estimation
Specia, Lucia, Raj, Dhwaj, and Turchi, Marco · 2010
Earlier work this paper cites.
The Hidden Information State model: A practical framework for POMDP-based spoken dialogue management
Young, Steve, Gašić, Milica, Keizer, Simon, Mairesse, François, Schatzmann, Jost, Thomson, Blaise, and Yu, Kai · 2010
Earlier work this paper cites.
Findings of the 2012 Workshop on Statistical Machine Translation
Callison-Burch, Chris, Koehn, Philipp, Monz, Christof, Post, Matt, Soricut, Radu, and Specia, Lucia · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, Geoffrey E., Srivastava, Nitish, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan R · 2012
Earlier work this paper cites.
The UI System in the HOO 2012 Shared Task on Error Correction
Rozovskaya, Alla, Sammons, Mark, and Roth, Dan · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Bojar, Ondřej, Buck, Christian, Callison-Burch, Chris, Federmann, Christian, Haddow, Barry, Koehn, Philipp, Monz, Christof, Post, Matt, Soricut, Radu, and Specia, Lucia · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Mikolov, Tomas, Chen, Kai, Corrado, Greg, and Dean, Jeffrey · 2013
Cited alongside, same era.
Findings of the 2014 Workshop on Statistical Machine Translation
Bojar, Ondřej, Buck, Christian, Federmann, Christian, Haddow, Barry, Koehn, Philipp, Leveling, Johannes, Monz, Christof, Pecina, Pavel, Post, Matt, Saint-Amand, Herve, Soricut, Radu, Specia, Lucia, and Tamchyna, Aleš · 2014
Cited alongside, same era.
A systematic comparison of smoothing techniques for sentence-level BLEU
Chen, Boxing and Cherry, Colin · 2014
Cited alongside, same era.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, Kyunghyun, van Merrienboer, Bart, Gulcehre, Caglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
CIDEr: Consensus-Based Image Description Evaluation
Vedantam, Ramakrishna, Lawrence Zitnick, C., and Parikh, Devi · 2015
Later among the works it cites.
Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
Wen, Tsung-Hsien, Gasic, Milica, Mrkšić, Nikola, Su, Pei-Hao, Vandyke, David, and Young, Steve · 2015
Later among the works it cites.
Findings of the 2016 conference on machine translation (WMT16)
Bojar, Ondrej, Chatterjee, Rajen, Federmann, Christian, Graham, Yvette, Haddow, Barry, Huck, Matthias, Yepes, Antonio Jimeno, Koehn, Philipp, Logacheva, Varvara, Monz, Christof, and others · 2016
Later among the works it cites.
Recurrent Neural Network based Translation Quality Estimation
Kim, Hyun and Lee, Jong-Hyeok · 2016
Later among the works it cites.
Imitation learning for language generation from unaligned data
Lampouras, Gerasimos and Vlachos, Andreas · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cluster-based prediction of user ratings for stylistic surface realisation
Dethlefs, Nina, Cuayáhuitl, Heriberto, Hastie, Helen F, Rieser, Verena, and Lemon, Oliver · 2014
Cited alongside, same era.
Generating artificial errors for grammatical error correction
Felice, Mariano and Yuan, Zheng · 2014
Cited alongside, same era.
Robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation
Henderson, M., Thomson, B., and Young, S · 2014
Cited alongside, same era.
Findings of the 2015 Workshop on Statistical Machine Translation
Bojar, Ondřej, Chatterjee, Rajen, Federmann, Christian, Haddow, Barry, Huck, Matthias, Hokamp, Chris, Koehn, Philipp, Logacheva, Varvara, Monz, Christof, Negri, Matteo, Post, Matt, Scarton, Carolina, Specia, Lucia, and Turchi, Marco · 2015
Cited alongside, same era.
Microsoft COCO Captions: Data Collection and Evaluation Server
Chen, Xinlei, Fang, Hao, Lin, Tsung-Yi, Vedantam, Ramakrishna, Gupta, Saurabh, Dollar, Piotr, and Zitnick, C. Lawrence · 2015
Cited alongside, same era.
Training a Natural Language Generator From Unaligned Data
Dušek, Ondřej and Jurčíček, Filip · 2015
Cited alongside, same era.
A Snapshot of NLG Evaluation Practices 2005 - 2014
Gkatzia, Dimitra and Mahamood, Saad · 2015
Cited alongside, same era.
Liu, Chia-Wei, Lowe, Ryan, Serban, Iulian, Noseworthy, Michael, Charlin, Laurent, and Pineau, Joelle · 2016
Later among the works it cites.
There’s No Comparison: Reference-less Evaluation Metrics in Grammatical Error Correction
Napoles, Courtney, Sakaguchi, Keisuke, and Tetreault, Joel · 2016
Later among the works it cites.
Translation Quality Estimation using Recurrent Neural Network
Patel, Raj Nath and M, Sasikumar · 2016
Later among the works it cites.
Neural Language Correction with Character-Based Attention
Xie, Ziang, Avati, Anand, Arivazhagan, Naveen, Jurafsky, Dan, and Ng, Andrew Y · 2016
Later among the works it cites.
A review of evaluation techniques for social dialogue systems
Cercas Curry, Amanda, Hastie, Helen, and Rieser, Verena · 2017
Closest in time.
Re-evaluating Automatic Metrics for Image Captioning
Kilickaya, Mert, Erdem, Aykut, Ikizler-Cinbis, Nazli, and Erdem, Erkut · 2017
Closest in time.
Towards an automatic Turing test: learning to evaluate dialogue responses
Lowe, Ryan, Noseworthy, Michael, Serban, Iulian V., Angelard-Gontier, Nicolas, Bengio, Yoshua, and Pineau, Joelle · 2017
Closest in time.
Why we need new evaluation metrics for NLG
Novikova, Jekaterina, Cercas Curry, Amanda, Dušek, Ondřej, and Rieser, Verena · 2017
Closest in time.