Fetching the paper…
Reading the bibliography…
We perform an in-depth error analysis of Adversarial NLI (ANLI), a recently introduced large-scale human-and-model-in-the-loop natural language inference dataset collected over multiple rounds.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 1908
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Probing natural language inference models through semantic fragments
Kyle Richardson, Hai Hu, Lawrence S Moss, and Ashish Sabharwal. 2019 · 1909
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Universal grammar
Richard Montague. 1970 · 1970
Earlier work this paper cites.
Logic and conversation
H Paul Grice. 1975 · 1975
Earlier work this paper cites.
An application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers
J. Richard Landis and Gary G. Koch. 1977 · 1977
Earlier work this paper cites.
Using the framework. technical report lre 62-051r
Robin Cooper, Crouch Dick, Jan van Eijck, Chris Fox, Joseph van Genabith, Han Jaspars, Hans Kamp, David Milward, Manfred Pinkal, Massimo Poesio, Steve Pulman, Ted Brisco, Holger Maier, and Karsten Konrad. 1996 · 1996
Earlier work this paper cites.
Collecting entailment data for pretraining: New protocols and negative results
Samuel R Bowman, Jennimaria Palomaki, Livio Baldini Soares, and Emily Pitler. 2020 · 2004
Earlier work this paper cites.
Scalar implicatures, polarity phenomena, and the syntax/pragmatics interface
Gennaro Chierchia et al. 2004 · 2004
Earlier work this paper cites.
The curse of performance instability in analysis datasets: Consequences, source, and suggestions
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020 · 2004
Earlier work this paper cites.
VerbNet: A broad-coverage, comprehensive verb lexicon
Karin Kipper Schuler. 2005 · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
Content analysis: What are they talking about?
Jan-Willem Strijbos, Rob L. Martens, Frans J. Prins, and Wim M.G. Jochems. 2006 · 2006
Earlier work this paper cites.
Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008 · 2008
Earlier work this paper cites.
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020b · 2010
Earlier work this paper cites.
Conjnli: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal. 2020 · 2010
Earlier work this paper cites.
“ask not what textual entailment can do for you…”
Mark Sammons, V.G.Vinod Vydiswaran, and Dan Roth. 2010 · 2010
Earlier work this paper cites.
Types of common-sense knowledge needed for recognizing textual entailment
Peter LoBue and Alexander Yates. 2011 · 2011
Earlier work this paper cites.
Developing a large semantically annotated corpus
Valerio Basile, Johan Bos, Kilian Evang, and Noortje Venhuizen. 2012 · 2012
Cited alongside, same era.
Cognitive and psychometric analysis of analogical problem solving
Isaac I Bejar, Roger Chaffin, and Susan Embretson. 2012 · 2012
Cited alongside, same era.
SemEval-2012 task 2: Measuring degrees of relational similarity
David Jurgens, Saif Mohammad, Peter Turney, and Keith Holyoak. 2012 · 2012
Cited alongside, same era.
Interrater reliability: the kappa statistic
Mary L. McHugh. 2012 · 2012
Cited alongside, same era.
Semantic annotation for textual entailment recognition
Assaf Toledo, Sophia Katrenko, Stavroula Alexandropoulou, Heidi Klockmann, Asher Stern, Ido Dagan, and Yoad Winter. 2012 · 2012
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Corpus of linguistic acceptability
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
The role of veridicality and factivity in clause selection
Aaron Steven White and Kyle Rawlins. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
What kind of natural language inference are nlp systems learning: Is this enough?
Jean-Philippe Bernardy and Stergios Chatzikyriakidis. 2019 · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modeling natural language semantics in learned representations
Samuel R. Bowman. 2016 · 2016
Cited alongside, same era.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
The Groningen Meaning Bank
Johan Bos, Valerio Basile, Kilian Evang, Noortje J Venhuizen, and Johannes Bjerva. 2017 · 2017
Cited alongside, same era.
Inference is everything: Recasting semantic resources into a unified evaluation framework
Aaron Steven White, Pushpendre Rastogi, Kevin Duh, and Benjamin Van Durme. 2017 · 2017
Cited alongside, same era.
What knowledge is needed to solve the RTE5 textual entailment challenge?
Peter Clark. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Human vs. muppet: A conservative estimate of human performance on the GLUE benchmark
Nikita Nangia and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Analyzing compositionality-sensitivity of NLI models
Yixin Nie, Yicheng Wang, and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
EQUATE: A benchmark evaluation framework for quantitative reasoning in natural language inference
Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy. 2019 · 2019
Later among the works it cites.
Diversify your datasets: Analyzing generalization via controlled variance in adversarial datasets
Ohad Rozen, Vered Shwartz, Roee Aharoni, and Ido Dagan. 2019 · 2019
Later among the works it cites.
Investigating BERT’s knowledge of language: Five analysis methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
HELP: A dataset for identifying shortcomings of neural models in monotonicity reasoning
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen-tau Yih, and Yejin Choi. 2020 · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Uncertain natural language inference
Tongfei Chen, Zhengping Jiang, Adam Poliak, Keisuke Sakaguchi, and Benjamin Van Durme. 2020 · 2020
Closest in time.
Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020 · 2020
Closest in time.
AmbigQA: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.