Fetching the paper…
Reading the bibliography…
Recognizing Textual Entailment (RTE) was proposed as a unified evaluation framework to compare semantic understanding of different NLP systems.
Question answering is a format; when is it useful?
Matt Gardner, Jonathan Berant, Hannaneh Hajishirzi, Alon Talmor, and Sewon Min. 2019 · 1909
Earlier work this paper cites.
Probing natural language inference models through semantic fragments
Kyle Richardson, Hai Na Hu, Lawrence S. Moss, and Ashish Sabharwal. 2020 · 1909
Earlier work this paper cites.
A deductive question-answerer for natural language inference
Robert M Schwarcz, John F Burger, and Robert F Simmons. 1970 · 1970
Earlier work this paper cites.
A preferential, pattern-seeking, semantics for natural language inference
Yorick Wilks. 1975 · 1975
Earlier work this paper cites.
Using test suites in evaluation of machine translation systems
Margaret King and Kirsten Falkedal. 1990 · 1990
Earlier work this paper cites.
Workshop on the evaluation of natural language processing systems
Martha Palmer and Tim Finin. 1990 · 1990
Earlier work this paper cites.
MUC-3 linguistic phenomena test experiment
Nancy Chinchor. 1991 · 1991
Earlier work this paper cites.
Evaluating message understanding systems: An analysis of the third message understanding conference (MUC-3)
Nancy Chinchor, Lynette Hirschman, and David D. Lewis. 1993 · 1993
Earlier work this paper cites.
Towards better NLP system evaluation
Karen Sparck Jones. 1994 · 1994
Earlier work this paper cites.
Eagles: Evaluation of natural language processing systems. final report
Maghi King, Bente MAEGAARD, Jorg SCHÜTZ, Louis des TOMBE, Annelise BECH, Ann NEVILLE, Antti ARPPE, Lorna BALKAN, Colin BRACE, Harry BUNT, Lauri CARLSON, Shona DOUGLAS, Monika HÖGE, Steven KRAUWER, Sandra MANZI, Cristina MAZZI, Ann June SIELEMANN, and Ragna STEENBAKKERS. 1995 · 1995
Earlier work this paper cites.
Tsnlp - test suites for natural language processing
Stephan Oepen and Klaus Netter. 1995 · 1995
Earlier work this paper cites.
Using the framework
Robin Cooper, Dick Crouch, Jan Van Eijck, Chris Fox, Johan Van Genabith, Jan Jaspars, Hans Kamp, David Milward, Manfred Pinkal, Massimo Poesio, et al. 1996 · 1996
Earlier work this paper cites.
TSNLP - test suites for natural language processing
Sabine Lehmann, Stephan Oepen, Sylvie Regnier-Prost, Klaus Netter, Veronika Lux, Judith Klein, Kirsten Falkedal, Frederik Fouvry, Dominique Estival, Eva Dauphin, Herve Compagnion, Judith Baur, Lorna Balkan, and Doug Arnold. 1996 · 1996
Earlier work this paper cites.
Evaluating Natural Language Processing Systems: An Analysis and Review
Karen Sparck Jones and Julia R. Galliers. 1996 · 1996
Earlier work this paper cites.
Karen sparck jones & julia r. galliers, evaluating natural language processing systems: An analysis and review. lecture notes in artificial intelligence 1083
Dominique Estival. 1997 · 1997
Earlier work this paper cites.
Senseval: an exercise in evaluating world sense disambiguation programs
Adam Kilgarriff. 1998 · 1998
Earlier work this paper cites.
Western Linguistics: An Historical Introduction
P.A.M. Seuren. 1998 · 1998
Earlier work this paper cites.
The Structure of Modern English: A linguistic introduction
L. Brinton. 2000 · 2000
Earlier work this paper cites.
Meaning and grammar: An introduction to semantics
Gennaro Chierchia and Sally McConnell-Ginet. 2000 · 2000
Earlier work this paper cites.
Workshop on MT Evaluation: Hands-On Evaluation
2001 · 2001
Earlier work this paper cites.
Placing search in context: The concept revisited
Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2001 · 2001
Earlier work this paper cites.
A test suite for evaluation of english-to-korean machine translation systems
Sungryong Koh, Jinee Maeng, Ji-Young Lee, Young-Sook Chae, and Key-Sun Choi. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Intrinsic versus extrinsic evaluations of parsing systems
Diego Mollá and Ben Hutchinson. 2003 · 2003
Earlier work this paper cites.
Proceedings of the EACL 2003 Workshop on Evaluation Initiatives in Natural Language Processing: are evaluation methods, metrics and resources reusable? Association for Computational Linguistics, Columbus, Ohio
Katerina Pastra, editor. 2003 · 2003
Earlier work this paper cites.
Letsum, an automatic legal text summarizing system
Atefeh Farzindar and Guy Lapalme. 2004 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Natural language inference via dependency tree mapping: An application to question answering
Vasin Punyakanok, Dan Roth, and Wen-tau Yih. 2004 · 2004
Earlier work this paper cites.
Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization . Association for Computational Linguistics, Ann Arbor, Michigan
Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss, editors. 2005 · 2005
Earlier work this paper cites.
Local textual inference: Can it be defined or circumscribed?
Annie Zaenen, Lauri Karttunen, and Richard Crouch. 2005 · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
Applied textual entailment
Oren Glickman. 2006 · 2006
Earlier work this paper cites.
Local textual inference: it’s hard to circumscribe, but you know it when you see it–and nlp needs it
Christopher D Manning. 2006 · 2006
Earlier work this paper cites.
Using intrinsic and extrinsic metrics to evaluate accuracy and facilitation in computer-assisted coding
Philip Resnik, Michael Niv, Michael Nossal, and Gregory Schnitzer. 2006 · 2006
Earlier work this paper cites.
What syntax can contribute in the entailment task
Lucy Vanderwende and William B Dolan. 2006 · 2006
Earlier work this paper cites.
The future of large-scale evaluation campaigns for information retrieval in europe
Maristella Agosti, Giorgio Maria Di Nunzio, Nicola Ferro, Donna Harman, and Carol Peters. 2007 · 2007
Earlier work this paper cites.
The role of sentence structure in recognizing textual entailment
Catherine Blake. 2007 · 2007
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
Principles of evaluation in natural language processing
Patrick Paroubek, Stéphane Chaudiron, and Lynette Hirschman. 2007 · 2007
Earlier work this paper cites.
Linguistic-Based Computational Treatment of Textual Entailment Recognition
Marilisa Amoia. 2008 · 2008
Earlier work this paper cites.
Intrinsic vs. extrinsic evaluation measures for referring expression generation
Anja Belz and Albert Gatt. 2008 · 2008
Earlier work this paper cites.
Intrinsic evaluation of text mining tools may not predict performance on realistic tasks
J Gregory Caporaso, Nita Deshpande, J Lynn Fink, Philip E Bourne, K Bretonnel Cohen, and Lawrence Hunter. 2008 · 2008
Earlier work this paper cites.
Finding contradictions in text
Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning. 2008 · 2008
Cited alongside, same era.
Natural language inference
Bill MacCartney. 2009 · 2009
Cited alongside, same era.
Automatic evaluation of linguistic quality in multi-document summarization
Emily Pitler, Annie Louis, and Ani Nenkova. 2010 · 2010
Cited alongside, same era.
11 evaluation of nlp systems
Philip Resnik and Jimmy Lin. 2010 · 2010
Cited alongside, same era.
Montague semantics
Theo M. V. Janssen. 2011 · 2011
Cited alongside, same era.
Distributional semantics in technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, and Nam-Khanh Tran. 2012 · 2012
Cited alongside, same era.
Breaking nli systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Later among the works it cites.
Natural language inference over interaction space
Yichen Gong, Heng Luo, and Jian Zhang. 2018 · 2018
Later among the works it cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Later among the works it cites.
Visualisation and’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Later among the works it cites.
SciTail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proceedings of Workshop on Evaluation Metrics and System Comparison for Automatic Summarization . Association for Computational Linguistics, Montréal, Canada
John M. Conroy, Hoa Trang Dang, Ani Nenkova, and Karolina Owczarzak, editors. 2012 · 2012
Cited alongside, same era.
Semeval-2013 task 7: The joint student response analysis and 8th recognizing textual entailment challenge
Myroslava Dzikovska, Rodney Nielsen, Chris Brew, Claudia Leacock, Danilo Giampiccolo, Luisa Bentivogli, Peter Clark, Ido Dagan, and Hoa Trang Dang. 2013 · 2013
Cited alongside, same era.
Better word representations with recursive neural networks for morphology
Thang Luong, Richard Socher, and Christopher Manning. 2013 · 2013
Cited alongside, same era.
An unsupervised model for instance level subcategorization acquisition
Simon Baker, Roi Reichart, and Anna Korhonen. 2014 · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP . Association for Computational Linguistics, Brussels, Belgium
Tal Linzen, Grzegorz Chrupała, and Afra Alishahi, editors. 2018 · 2018
Later among the works it cites.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018 · 2018
Later among the works it cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Later among the works it cites.
Word embeddings in sentiment analysis
Ruggero Petrolito. 2018 · 2018
Later among the works it cites.
On the evaluation of semantic phenomena in neural machine translation using natural language inference
Adam Poliak, Yonatan Belinkov, James Glass, and Benjamin Van Durme. 2018a · 2018
Later among the works it cites.
Collecting diverse natural language inference problems for sentence representation evaluation
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018b · 2018
Later among the works it cites.
Revisiting correlations between intrinsic and extrinsic evaluations of word embeddings
Yuanyuan Qiu, Hongzheng Li, Shen Li, Yingdi Jiang, Renfen Hu, and Lijiao Yang. 2018 · 2018
Later among the works it cites.
A structured review of the validity of bleu
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Lessons from natural language inference in the clinical domain
Alexey Romanov and Chaitanya Shivade. 2018 · 2018
Later among the works it cites.
Learning about non-veridicality in textual entailment
Ieva Staliūnaitė. 2018 · 2018
Later among the works it cites.
Compare, compress and propagate: Enhancing neural architectures with alignment factorization for natural language inference
Yi Tay, Anh Tuan Luu, and Siu Cheung Hui. 2018 · 2018
Later among the works it cites.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Masatoshi Tsuchiya. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Linguistic evaluation of German-English machine translation using a test suite
Eleftherios Avramidis, Vivien Macketanz, Ursula Strohriegel, and Hans Uszkoreit. 2019 · 2019
Later among the works it cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Later among the works it cites.
Automatically extracting challenge sets for non-local phenomena in neural machine translation
Leshem Choshen and Omri Abend. 2019 · 2019
Later among the works it cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Probing what different NLP tasks teach machines about function word comprehension
Najoung Kim, Roma Patel, Adam Poliak, Patrick Xia, Alex Wang, Tom McCoy, Ian Tenney, Alexis Ross, Tal Linzen, Benjamin Van Durme, Samuel R. Bowman, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Temporal and aspectual entailment
Thomas Kober, Sander Bijl de Vroe, and Mark Steedman. 2019 · 2019
Later among the works it cites.
Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP . Association for Computational Linguistics, Florence, Italy
Tal Linzen, Grzegorz Chrupała, Yonatan Belinkov, and Dieuwke Hupkes, editors. 2019 · 2019
Later among the works it cites.
Performance evaluation of word and sentence embeddings for finance headlines sentiment analysis
Kostadin Mishev, Ana Gjorgjevikj, Riste Stojanov, Igor Mishkovski, Irena Vodenska, Ljubomir Chitkushev, and Dimitar Trajanov. 2019 · 2019
Later among the works it cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
Challenge test sets for MT evaluation
Maja Popović and Sheila Castilho. 2019 · 2019
Later among the works it cites.
Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP . Association for Computational Linguistics, Minneapolis, USA
Anna Rogers, Aleksandr Drozd, Anna Rumshisky, and Yoav Goldberg, editors. 2019 · 2019
Later among the works it cites.
How well do NLI models capture verb veridicality?
Alexis Ross and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
SWOW-8500: Word association task for intrinsic evaluation of word embeddings
Avijit Thawani, Biplav Srivastava, and Anil Singh. 2019 · 2019
Later among the works it cites.
Can neural networks understand monotonicity reasoning?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 · 2019
Later among the works it cites.
Adversarial nli for factual correctness in text summarisation models
Mario Barrantes, Benedikt Herudek, and Richard Wang. 2020 · 2020
Closest in time.
Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020 · 2020
Closest in time.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas. 2020 · 2020
Closest in time.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Closest in time.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Closest in time.
Temporal reasoning in natural language inference
Siddharth Vashishtha, Adam Poliak, Yash Kumar Lal, Benjamin Van Durme, and Aaron Steven White. 2020 · 2020
Closest in time.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Closest in time.
Do neural models learn systematicity of monotonicity inference in natural language?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, and Kentaro Inui. 2020 · 2020
Closest in time.
Evaluation of word vector representations by subspace alignment
Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Guillaume Lample, and Chris Dyer. 2015 · 2054
Closest in time.