Fetching the paper…
Reading the bibliography…
There has always been criticism for using $n$-gram based similarity metrics, such as BLEU, NIST, etc, for evaluating the performance of NLG systems.
Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit
J Cohen. 1968 · 1968
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017 · 1997
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Confidence estimation for machine translation
John Blatz, Erin Fitzgerald, George Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza, Alberto Sanchis, and Nicola Ueffing. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Some issues in automatic evaluation of english-hindi mt: More blues for bleu
Ananthakrishnan R, Pushpak Bhattacharyya, M Sasikumar, and Ritesh M Shah. 2007 · 2007
Earlier work this paper cites.
Fast, cheap, and creative: Evaluating translation quality using amazon’s mechanical turk
Chris Callison-Burch. 2009 · 2009
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael J. Denkowski. 2009 · 2009
Earlier work this paper cites.
Generating instruction automatically for the reading strategy of self-questioning
Jack Mostow and Wei Chen. 2009 · 2009
Earlier work this paper cites.
Good question! statistical ranking for question generation
Michael Heilman and Noah A. Smith. 2010 · 2010
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh. 2012 · 2012
Cited alongside, same era.
Generating natural language questions to support learning on-line
David Lindberg, Fred Popowich, John C. Nesbit, and Philip H. Winne. 2013 · 2013
Cited alongside, same era.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Large-scale simple question answering with memory networks
Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Deep questions without deep understanding
Igor Labutov, Sumit Basu, and Lucy Vanderwende. 2015 · 2015
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Question generation for question answering
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017 · 2017
Later among the works it cites.
Creativity: Generating diverse questions using variational autoencoders
Unnat Jain, Ziyu Zhang, and Alexander G. Schwing. 2017 · 2017
Later among the works it cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Later among the works it cites.
Visual question generation as dual task of visual question answering
Yikang Li, Nan Duan, Bolei Zhou, Xiao Chu, Wanli Ouyang, and Xiaogang Wang. 2017 · 2017
Later among the works it cites.
Towards an automatic turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Alexander H. Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016 · 2016
Cited alongside, same era.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Cited alongside, same era.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Bidirectional attention flow for machine comprehension
Min Joon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016 · 2016
Cited alongside, same era.
Generating factoid questions with recurrent neural networks: The 30m factoid question-answer corpus
Iulian Vlad Serban, Alberto García-Durán, Çaglar Gülçehre, Sungjin Ahn, Sarath Chandar, Aaron C. Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Cited alongside, same era.
Generating natural language question-answer pairs from a knowledge graph using a RNN based question generation model
Sathish Reddy, Dinesh Raghu, Mitesh M. Khapra, and Sachindra Josh. 2017 · 2017
Later among the works it cites.
Question answering and question generation as dual tasks
Duyu Tang, Nan Duan, Tao Qin, and Ming Zhou. 2017 · 2017
Later among the works it cites.
A joint model for question answering and question generation
Tong Wang, Xingdi Yuan, and Adam Trischler. 2017 · 2017
Later among the works it cites.
Semi-supervised QA with generative domain-adaptive nets
Zhilin Yang, Junjie Hu, Ruslan Salakhutdinov, and William W. Cohen. 2017 · 2017
Later among the works it cites.
Machine comprehension by text-to-text neural question generation
Xingdi Yuan, Tong Wang, Çaglar Gülçehre, Alessandro Sordoni, Philip Bachman, Saizheng Zhang, Sandeep Subramanian, and Adam Trischler. 2017 · 2017
Later among the works it cites.
Asking the difficult questions: Goal-oriented visual question generation via intermediate rewards
Junjie Zhang, Qi Wu, Chunhua Shen, Jian Zhang, Jianfeng Lu, and Anton van den Hengel. 2017 · 2017
Later among the works it cites.
Identifying where to focus in reading comprehension for neural question generation
Xinya Du and Claire Cardie. 2017 · 2073
Closest in time.