Fetching the paper…
Reading the bibliography…
Posing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
On the evaluation of machine translation systems trained with back-translation
Sergey Edunov, Myle Ott, Marc’Aurelio Ranzato, and Michael Auli. 2019 · 1908
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, S. Gross, Francisco Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, Alban Desmaison, Andreas Köpf, E. Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, B. Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 1912
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2001 · 2001
Earlier work this paper cites.
Evaluating nlp models via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Quan Zhang, and Ben Zhou. 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
A. R. Fabbri, Wojciech Kryscinski, B. McCann, R. Socher, and D. Radev. 2020 · 2007
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
K. Krippendorff. 2011 · 2011
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
Results of the wmt14 metrics shared task
Matous Machácek and Ondrej Bojar. 2014 · 2014
Earlier work this paper cites.
chrf: character n-gram f-score for automatic mt evaluation
Maja Popovic. 2015 · 2015
Earlier work this paper cites.
Results of the wmt15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondrej Bojar. 2015 · 2015
Earlier work this paper cites.
Results of the wmt16 metrics shared task
Ondrej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Results of the wmt17 metrics shared task
Ondrej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Matthew Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson H S Liu, Matthew E. Peters, Michael Schmitz, and Luke Zettlemoyer. 2017 · 2017
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Later among the works it cites.
Evaluating question answering evaluation
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, and Matt Gardner. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Cited alongside, same era.
The narrativeqa reading comprehension challenge
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2017 · 2017
Cited alongside, same era.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. 2017 · 2017
Cited alongside, same era.
Blend: a novel combined mt metric based on direct assessment - casict-dcu submission to wmt17 metrics task
Qingsong Ma, Yvette Graham, Shugen Wang, and Qun Liu. 2017 · 2017
Cited alongside, same era.
Commonsense for generative multi-hop question answering tasks
Lisa Bauer, Yicheng Wang, and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
Learning to evaluate image captioning
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge J. Belongie. 2018 · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Cosmos qa: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Results of the wmt19 metrics shared task: Segment-level and strong mt systems pose big challenges
Qingsong Ma, Johnny Wei, Ondrej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
Compositional questions do not necessitate multi-hop reasoning
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Social iqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020 · 2020
Closest in time.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh. 2020 · 2020
Closest in time.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Michael Lewis. 2020 · 2020
Closest in time.