Fetching the paper…
Reading the bibliography…
In the age of large transformer language models, linguistic evaluation play an important role in diagnosing models' abilities and limitations on natural language understanding.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
A comparison of mcc and cen error measures in multi-class prediction
Giuseppe Jurman, Samantha Riccadonna, and Cesare Furlanello. 2012 · 2012
Earlier work this paper cites.
Recognizing Textual Entailment: Models and Applications
Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Inference is everything: Recasting semantic resources into a unified evaluation framework
Aaron Steven White, Pushpendre Rastogi, Kevin Duh, and Benjamin Van Durme. 2017 · 2017
Earlier work this paper cites.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Earlier work this paper cites.
Collecting diverse natural language inference problems for sentence representation evaluation
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018a · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
ATOMIC: an atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan LeBras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Verb argument structure alternations in word and sentence embeddings
Katharina Kann, Alex Warstadt, Adina Williams, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Inoculation by fine-tuning: A method for analyzing challenge datasets
Nelson F. Liu, Roy Schwartz, and Noah A. Smith. 2019a · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Probing natural language inference models through semantic fragments
Kyle Richardson, Hai Hu, Lawrence S. Moss, and Ashish Sabharwal. 2019 · 2019
Cited alongside, same era.
Diversify your datasets: Analyzing generalization via controlled variance in adversarial datasets
Ohad Rozen, Vered Shwartz, Roee Aharoni, and Ido Dagan. 2019 · 2019
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
What does my QA model know? devising controlled probes using expert knowledge
Kyle Richardson and Ashish Sabharwal. 2020 · 2020
Later among the works it cites.
Thinking like a skeptic: Defeasible inference in natural language
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Beyond the imitation game: Measuring and extrapolating the capabilities of language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019 · 2019
Cited alongside, same era.
Can neural networks understand monotonicity reasoning?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019a · 2019
Cited alongside, same era.
HELP: A dataset for identifying shortcomings of neural models in monotonicity reasoning
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019b · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
BIG-bench collaboration. 2021 · 2021
Later among the works it cites.
Probing linguistic information for logical inference in pre-trained language models
Zeming Chen and Qiyue Gao. 2021 · 2021
Later among the works it cites.
Explaining answers with entailment trees
Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021 · 2021
Later among the works it cites.
Information-theoretic measures of dataset difficulty
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2021 · 2021
Later among the works it cites.
Ester: A machine reading comprehension dataset for event semantic relation reasoning
Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon, Qiang Ning, Dan Roth, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
Can transformer language models predict psychometric properties?
Antonio Laverghetta Jr., Animesh Nighojkar, Jamshidbek Mirzakhalov, and John Licato. 2021 · 2021
Later among the works it cites.
Natural language inference in context - investigating contextual reasoning over long texts
Hanmeng Liu, Leyang Cui, Jian Liu, and Yue Zhang. 2021 · 2021
Later among the works it cites.
Ai and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Language models for lexical inference in context
Martin Schmitt and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Lonli: An extensible framework for testing diverse logical reasoning capabilities for nli
Ishan Tarunesh, Somak Aditya, and Monojit Choudhury. 2021 · 2021
Later among the works it cites.
Exploring transitivity in neural NLI models through veridicality
Hitomi Yanaka, Koji Mineshima, and Kentaro Inui. 2021 · 2021
Later among the works it cites.
Ar-lsat: Investigating analytical reasoning of text
Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan. 2021 · 2021
Later among the works it cites.
Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts
Zhengxiang Shi, Qiang Zhang, and Aldo Lipani. 2022 · 2022
Closest in time.