Fetching the paper…
Reading the bibliography…
As major progress is made in open-ended text generation, measuring how close machine-generated text is to human language remains a critical open problem.
A learning algorithm for Boltzmann machines
D. H. Ackley, G. E. Hinton, and T. J. Sejnowski · 1985
Earlier work this paper cites.
Analyzing and modeling rank data , volume 64 of Monographs on Statistics and Applied Probability
J. I. Marden · 1995
Earlier work this paper cites.
Foundations of Statistical Natural Language Processing
C. D. Manning and H. Schütze · 2001
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
MM algorithms for generalized Bradley-Terry models
D. R. Hunter · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Confidence intervals for the area under the ROC curve
C. Cortes and M. Mohri · 2005
Earlier work this paper cites.
Mathematics of Information and Coding , volume 203
T. S. Han and K. Kobayashi · 2007
Earlier work this paper cites.
Nonparametric estimation of the precision-recall curve
S. Clémençon and N. Vayatis · 2009
Earlier work this paper cites.
Overlaying classifiers: a practical approach to optimal scoring
S. Clémençon and N. Vayatis · 2010
Earlier work this paper cites.
Machine Learning: The Art and Science of Algorithms That Make Sense of Data
P. Flach · 2012
Earlier work this paper cites.
Nonlinear Multiobjective Optimization , volume 12
K. Miettinen · 2012
Earlier work this paper cites.
Generative Adversarial Networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
From Word Embeddings to Document Distances
M. Kusner, Y. Sun, N. Kolkin, and K. Weinberger · 2015
Earlier work this paper cites.
From Softmax to Sparsemax: A Sparse model of Attention and Multi-label Classification
A. Martins and R. Astudillo · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs, 2016
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Revisiting Classifier Two-Sample Tests
D. Lopez-Paz and M. Oquab · 2017
Earlier work this paper cites.
Why We Need New Evaluation Metrics for NLG
J. Novikova, O. Dušek, A. Cercas Curry, and V. Rieser · 2017
Earlier work this paper cites.
Attention is All you Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Demystifying MMD GANs
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton · 2018
Cited alongside, same era.
Hierarchical Neural Story Generation
A. Fan, M. Lewis, and Y. N. Dauphin · 2018
Cited alongside, same era.
Assessing generative models via precision and recall
M. S. M. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly · 2018
Cited alongside, same era.
On Accurate Evaluation of GANs for Language Generation, 2018
S. Semeniuta, A. Severyn, and S. Gelly · 2018
Cited alongside, same era.
RUSE: Regressor Using Sentence Embeddingsfor Automatic Machine Translation Evaluation
H. Shimanaka, T. Kajiwara, and M. Komachi · 2018
Cited alongside, same era.
RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems
C. Tao, L. Mou, D. Zhao, and R. Yan · 2018
Cited alongside, same era.
Evaluation of Text Generation: A Survey
A. Celikyilmaz, E. Clark, and J. Gao · 2020
Later among the works it cites.
Precision-Recall Curves Using Information Divergence Frontiers
J. Djolonga, M. Lucic, M. Cuturi, O. Bachem, O. Bousquet, and S. Gelly · 2020
Later among the works it cites.
Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation
B. Eikema and W. Aziz · 2020
Later among the works it cites.
UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
J. Guan and M. Huang · 2020
Later among the works it cites.
Deep Residual Mixture Models
P. Hämäläinen and A. Solin · 2020
Later among the works it cites.
The Curious Case of Neural Text Degeneration
A. Holtzman, J. Buys, M. Forbes, and Y. Choi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Texygen: A Benchmarking Platform for Text Generation Models, 2018
Y. Zhu, S. Lu, L. Zheng, J. Guo, W. Zhang, J. Wang, and Y. Yu · 2018
Cited alongside, same era.
Sentence Mover’s Similarity: Automatic Evaluation for Multi-Sentence Texts
E. Clark, A. Celikyilmaz, and N. A. Smith · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
The Second Conversational Intelligence Challenge (ConvAI2), 2019
E. Dinan, V. Logacheva, V. Malykh, A. Miller, K. Shuster, J. Urbanek, D. Kiela, A. Szlam, I. Serban, R. Lowe, S. Prabhumoye, A. W. Black, A. Rudnicky, J. Williams, J. Pineau, M. Burtsev, and J. Weston · 2019
Cited alongside, same era.
Unifying human and statistical evaluation for natural language generation
T. Hashimoto, H. Zhang, and P. Liang · 2019
Cited alongside, same era.
Improved Precision and Recall Metric for Assessing Generative Models
T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila · 2019
Cited alongside, same era.
Automatic Detection of Generated Text is Easiest when Humans are Fooled
D. Ippolito, D. Duckworth, C. Callison-Burch, and D. Eck · 2020
Later among the works it cites.
Sparse Text Generation
P. H. Martins, Z. Marinho, and A. F. T. Martins · 2020
Later among the works it cites.
PlotTMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking
H. Rashkin, A. Celikyilmaz, Y. Choi, and J. Gao · 2020
Later among the works it cites.
A Survey of Evaluation Metrics Used for NLG Systems
A. B. Sai, A. K. Mohankumar, and M. M. Khapra · 2020
Later among the works it cites.
BLEURT: Learning Robust Metrics for Text Generation
T. Sellam, D. Das, and A. P. Parikh · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Later among the works it cites.
BERTScore: Evaluating text generation with BERT
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Later among the works it cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Closest in time.
All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text
E. Clark, T. August, S. Serrano, N. Haduong, S. Gururangan, and N. A. Smith · 2021
Closest in time.
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
M. Karpinska, N. Akoury, and M. Iyyer · 2021
Closest in time.
Divergence Frontiers for Generative Models: Sample Complexity, Quantization Effects, and Frontier Integrals
L. Liu, K. Pillutla, S. Welleck, S. Oh, Y. Choi, and Z. Harchaoui · 2021
Closest in time.
Towards a Decomposable Metric for Explainable Evaluation of Text Generation from AMR
J. Opitz and A. Frank · 2021
Closest in time.
The Human Evaluation Datasheet 1.0: A Template for Recording Details of Human Evaluation Experiments in NLP
A. Shimorina and A. Belz · 2021
Closest in time.
Topics in Optimal Transportation , volume 58
C. Villani · 2021
Closest in time.
Trading off diversity and quality in natural language generation
H. Zhang, D. Duckworth, D. Ippolito, and A. Neelakantan · 2021
Closest in time.