Fetching the paper…
Reading the bibliography…
We investigate how disagreement in natural language inference (NLI) annotation arises.
Bets and beliefs
Henry E. Kyburg. 1968 · 1968
Earlier work this paper cites.
Logic and conversation
Herbert P Grice. 1975 · 1975
Earlier work this paper cites.
Studies on the Semantics of Questions and the Pragmatics of Answers
Jeroen Antonius Gerardus Groenendijk and Martin Johan Bastiaan Stokhof. 1984 · 1984
Earlier work this paper cites.
A probabilistic setting and lexical coocurrence model for textual entailment
Oren Glickman and Ido Dagan. 2005 · 2005
Earlier work this paper cites.
The reliability of anaphoric annotation, reconsidered: Taking ambiguity into account
Massimo Poesio and Ron Artstein. 2005 · 2005
Earlier work this paper cites.
The Logic of Conventional Implicatures
Christopher Potts. 2005 · 2005
Earlier work this paper cites.
Measuring agreement on set-valued items (MASI) for semantic and pragmatic annotation
Rebecca Passonneau. 2006 · 2006
Earlier work this paper cites.
Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing . Association for Computational Linguistics, Prague
Satoshi Sekine, Kentaro Inui, Ido Dagan, Bill Dolan, Danilo Giampiccolo, and Bernardo Magnini, editors. 2007 · 2007
Earlier work this paper cites.
Finding contradictions in text
Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning. 2008 · 2008
Earlier work this paper cites.
Vagueness and Referential Ambiguity in a Large-Scale Annotated Corpus
Yannick Versley. 2008 · 2008
Earlier work this paper cites.
Graded word sense assignment
Katrin Erk and Diana McCarthy. 2009 · 2009
Earlier work this paper cites.
Assessing the role of discourse references in entailment inference
Shachar Mirkin, Ido Dagan, and Sebastian Padó. 2010 · 2010
Earlier work this paper cites.
“Ask not what textual entailment can do for you…”
Mark Sammons, V.G.Vinod Vydiswaran, and Dan Roth. 2010 · 2010
Earlier work this paper cites.
What projects and why
Mandy Simons, Judith Tonhauser, David Beaver, and Craige Roberts. 2010 · 2010
Earlier work this paper cites.
Types of common-sense knowledge needed for recognizing textual entailment
Peter LoBue and Alexander Yates. 2011 · 2011
Earlier work this paper cites.
Identity, non-identity, and near-identity: Addressing the complexity of coreference
Marta Recasens, Eduard Hovy, and M. Antònia Martí. 2011 · 2011
Earlier work this paper cites.
Did it happen? The pragmatic complexity of veridicality assessment
Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012 · 2012
Earlier work this paper cites.
Multiplicity and word sense: Evaluating and learning from multiply labeled word sense annotations
Rebecca J. Passonneau, Vikas Bhardwaj, Ansaf Salleb-Aouissi, and Nancy Ide. 2012 · 2012
Cited alongside, same era.
Information structure in discourse: Towards an integrated formal theory of pragmatics
Craige Roberts. 2012 · 2012
Cited alongside, same era.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Cited alongside, same era.
Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Modification
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
THOMAS: The hegemonic OSU morphological analyzer using seq2seq
Byung-Doh Oh, Pranav Maneriker, and Nanjiang Jiang. 2019 · 2019
Later among the works it cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
A crowdsourced corpus of multiple judgments and disagreement on anaphoric interpretation
Massimo Poesio, Jon Chamberlain, Silviu Paun, Juntao Yu, Alexandra Uma, and Udo Kruschwitz. 2019 · 2019
Later among the works it cites.
Evaluating semantic accuracy of data-to-text generation with natural language inference
Ondřej Dušek and Zdeněk Kasner. 2020 · 2020
Later among the works it cites.
What can we learn from collective human opinions on natural language inference data?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Louise McNally. 2016 · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
Soft label memorization-generalization for natural language inference
John P. Lalor, Hao Wu, and Hong Yu. 2017 · 2017
Cited alongside, same era.
Sentiment analysis: It’s complicated!
Kian Kenyon-Dean, Eisha Ahmed, Scott Fujimoto, Jeremy Georges-Filteau, Christopher Glasz, Barleen Kaur, Auguste Lalande, Shruti Bhanderi, Robert Belfer, Nirmal Kanagasabai, Roman Sarrazingendron, Rohit Verma, and Derek Ruths. 2018 · 2018
Cited alongside, same era.
QED: A fact verification system for the FEVER shared task
Jackson Luken, Nanjiang Jiang, and Marie-Catherine de Marneffe. 2018 · 2018
Cited alongside, same era.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Did they answer? subjective acts and intents in conversational discourse
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, and Katrin Erk. 2021 · 2021
Later among the works it cites.
Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning
Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein. 2021 · 2021
Later among the works it cites.
Learning from disagreement: A survey
Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
Learning with different amounts of annotation: From zero to many labels
Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021 · 2021
Later among the works it cites.
Identifying inherent disagreement in natural language inference
Xinliang Frederick Zhang and Marie-Catherine de Marneffe. 2021 · 2021
Later among the works it cites.
Distributed NLI: learning to predict human opinion distributions for language reasoning
Xiang Zhou, Yixin Nie, and Mohit Bansal. 2021 · 2021
Later among the works it cites.
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
Scaling and disagreements: Bias, noise, and ambiguity
Alexandra Uma, Dina Almanea, and Massimo Poesio. 2022 · 2022
Closest in time.
ANLIzing the adversarial natural language inference dataset
Adina Williams, Tristan Thrush, and Douwe Kiela. 2022 · 2022
Closest in time.