Fetching the paper…
Reading the bibliography…
How can we measure whether a natural language generation system produces both high quality and diverse outputs? Human evaluation captures quality but not diversity, as it does not catch models that simply plagiarize from the training set.
Minimax estimation of maximum mean discrepancy with radial kernels
I. Tolstikhin, B. K. Sriperumbudur, and B. Scholkopf. 2016 · 1938
Earlier work this paper cites.
Computing machinery and intelligence
A. M. Turing. 1950 · 1950
Earlier work this paper cites.
Relations between entropy and error probability
M. Feder and N. Merhav. 1994 · 1994
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu. 2002 · 2002
Earlier work this paper cites.
Looking for a few good metrics: ROUGE and its evaluation
C. Lin and M. Rey. 2004 · 2004
Earlier work this paper cites.
The wisdom of crowds: Why the many are smarter than the few and how collective wisdom shapes business, economies, societies, and nations
J. Surowiecki. 2004 · 2004
Earlier work this paper cites.
The pyramid method: Incorporating human content selection variation in summarization evaluation
A. Nenkova, R. J. Passonneau, and K. McKeown. 2007 · 2007
Earlier work this paper cites.
Novelty and diversity in information retrieval evaluation
C. L. A. Clarke, M. Kolla, G. V. Cormack, O. Vechtomova, A. Ashkan, S. Büttcher, and I. MacKinnon. 2008 · 2008
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
A. Lavie and M. Denkowski. 2009 · 2009
Earlier work this paper cites.
Diversity-aware evaluation for paraphrase patterns
H. Shima and T. Mitamura. 2011 · 2011
Earlier work this paper cites.
Joint learning of a dual SMT system for paraphrase generation
H. Sun and M. Zhou. 2012 · 2012
Earlier work this paper cites.
The good judgment project: A large scale test of different methods of combining expert predictions
L. Ungar, B. Mellors, V. Satopää, J. Baron, P. Tetlock, J. Ramos, and S. Swift. 2012 · 2012
Earlier work this paper cites.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll’ar, and C. L. Zitnick. 2014 · 2014
Earlier work this paper cites.
deltableu: A discriminative metric for generation tasks with intrinsically diverse targets
M. Galley, C. Brockett, A. Sordoni, Y. Ji, M. Auli, C. Quirk, M. Mitchell, J. Gao, and B. Dolan. 2015 · 2015
Cited alongside, same era.
A neural network approach to context-sensitive generation of conversational responses
A. Sordoni, M. Galley, M. Auli, C. Brockett, Y. Ji, M. Mitchell, J. Nie, J. Gao, and B. Dolan. 2015 · 2015
Cited alongside, same era.
A note on the evaluation of generative models
L. Theis, A. van den Oord, and M. Bethge. 2015 · 2015
Cited alongside, same era.
Generating sentences from a continuous space
S. R. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Jozefowicz, and S. Bengio. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu. 2016 · 2016
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. 2017 · 2017
Later among the works it cites.
Opennmt: Open-source toolkit for neural machine translation
G. Klein, Y. Kim, Y. Deng, J. Senellart, and A. M. Rush. 2017 · 2017
Later among the works it cites.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, A. Ritter, and D. Jurafsky. 2017 · 2017
Later among the works it cites.
Towards an automatic turing test: Learning to evaluate dialogue responses
R. Lowe, M. Noseworthy, I. V. Serban, N. Angelard-Gontier, Y. Bengio, and J. Pineau. 2017 · 2017
Later among the works it cites.
Why we need new evaluation metrics for NLG
J. Novikova, O. Dušek, A. C. Curry, and V. Rieser. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adversarial evaluation of dialogue models
A. Kannan and O. Vinyals. 2016 · 2016
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
J. Li, M. Galley, C. Brockett, J. Gao, and W. B. Dolan. 2016 · 2016
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
C. Liu, R. Lowe, I. V. Serban, M. Noseworthy, L. Charlin, and J. Pineau. 2016 · 2016
Cited alongside, same era.
A corpus and cloze evaluation for deeper understanding of commonsense stories
N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen. 2016 · 2016
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al. 2016 · 2016
Cited alongside, same era.
Writing stories with help from recurrent neural networks
M. Roemmele. 2016 · 2016
Cited alongside, same era.
Hypothesis testing for high-dimensional multinomials: A selective review
S. Balakrishnan and L. Wasserman. 2017 · 2017
Cited alongside, same era.
M. Caccia, L. Caccia, W. Fedus, H. Larochelle, J. Pineau, and L. Charlin. 2018 · 2018
Later among the works it cites.
The price of debiasing automatic metrics in natural language evaluation
A. Chaganty, S. Mussmann, and P. Liang. 2018 · 2018
Later among the works it cites.
Bottom-up abstractive summarization
S. Gehrmann, Y. Deng, and A. M. Rush. 2018 · 2018
Later among the works it cites.
A retrieve-and-edit framework for predicting structured outputs
T. Hashimoto, K. Guu, Y. Oren, and P. Liang. 2018 · 2018
Later among the works it cites.
Retrieval-based neural code generation
S. A. Hayati, R. Olivier, P. Avvaru, P. Yin, A. Tomasic, and G. Neubig. 2018 · 2018
Later among the works it cites.
Delete, retrieve, generate: A simple approach to sentiment and style transfer
J. Li, R. Jia, H. He, and P. Liang. 2018 · 2018
Later among the works it cites.
Skill rating for generative models
C. Olsson, S. Bhupatiraju, T. Brown, A. Odena, and I. Goodfellow. 2018 · 2018
Later among the works it cites.
Assessing generative models via precision and recall
M. S. M. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly. 2018 · 2018
Later among the works it cites.
Nonparametric density estimation under adversarial losses
S. Singh, A. Uppal, B. Li, C. Li, M. Zaheer, and B. Poczos. 2018 · 2018
Later among the works it cites.