Fetching the paper…
Reading the bibliography…
Despite the subjective nature of many NLP tasks, most NLU evaluations have focused on using the majority label with presumably high agreement as the ground truth.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler. 1951 · 1951
Earlier work this paper cites.
Behavior analysis of nli models: Uncovering the influence of three factors on robustness
Ivan Sanchez, Jeff Mitchell, and Sebastian Riedel. 2018 · 1985
Earlier work this paper cites.
Information theory and statistics
Solomon Kullback. 1997 · 1997
Earlier work this paper cites.
Key concepts in language and linguistics
Robert Lawrence Trask. 1999 · 1999
Earlier work this paper cites.
A new metric for probability distributions
Dominik Maria Endres and Johannes E Schindelin. 2003 · 2003
Earlier work this paper cites.
The reliability of anaphoric annotation, reconsidered: Taking ambiguity into account
Massimo Poesio and Ron Artstein. 2005 · 2005
Earlier work this paper cites.
Local textual inference: it’s hard to circumscribe, but you know it when you see it–and nlp needs it
Christopher D Manning. 2006 · 2006
Earlier work this paper cites.
Exploiting ‘subjective’annotations
Dennis Reidsma and Rieks op den Akker. 2008 · 2008
Earlier work this paper cites.
Vagueness and referential ambiguity in a large-scale annotated corpus
Yannick Versley. 2008 · 2008
Earlier work this paper cites.
Graded word sense assignment
Katrin Erk and Diana McCarthy. 2009 · 2009
Earlier work this paper cites.
Did it happen? the pragmatic complexity of veridicality assessment
Marie-Catherine De Marneffe, Christopher D Manning, and Christopher Potts. 2012 · 2012
Earlier work this paper cites.
Embracing ambiguity: A comparison of annotation methodologies for crowdsourcing word sense labels
David Jurgens. 2013 · 2013
Cited alongside, same era.
The chameleon-like nature of evaluative adjectives
Lauri Karttunen, Stanley Peters, Annie Zaenen, and Cleo Condoravdi. 2014 · 2014
Cited alongside, same era.
Learning part-of-speech taggers with inter-annotator agreement loss
Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014 · 2014
Cited alongside, same era.
Learning to parse with iaa-weighted loss
Héctor Martínez Alonso, Barbara Plank, Arne Skjærholt, and Anders Søgaard. 2015 · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
A crowdsourced frame disambiguation corpus with ambiguity
Anca Dumitrache, Lora Aroyo, and Chris Welty. 2019 · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Later among the works it cites.
Human vs. muppet: A conservative estimate of human performance on the glue benchmark
Nikita Nangia and Samuel R Bowman. 2019 · 2019
Later among the works it cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
A crowdsourced corpus of multiple judgments and disagreement on anaphoric interpretation
Massimo Poesio, Jon Chamberlain, Silviu Paun, Juntao Yu, Alexandra Uma, and Udo Kruschwitz. 2019 · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ParlAI: A dialog research software platform
A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017 · 2017
Cited alongside, same era.
Ordinal common-sense inference
Sheng Zhang, Rachel Rudinger, Kevin Duh, and Benjamin Van Durme. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Uncertain natural language inference
Tongfei Chen, Zhengping Jiang, Keisuke Sakaguchi, and Benjamin Van Durme. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019 · 2019
Later among the works it cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2020 · 2020
Closest in time.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Closest in time.