Fetching the paper…
Reading the bibliography…
We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial.
FACT , pages 143–173. De Gruyter Mouton, Berlin, Boston
Paul Kiparsky and Carol Kiparsky. 1970 · 1970
Earlier work this paper cites.
Ordered Entailments: An Alternative to Presuppositional Theories , volume 11, pages 299–323
Deirdre Wilson and Dan Sperber. 1979 · 1979
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Adverb Licensing and Clause Structure in English
Dagmar Haumann. 2007 · 2007
Earlier work this paper cites.
What is presupposition accommodation, again?
Kai Fintel. 2008 · 2008
Earlier work this paper cites.
Modeling semantic containment and exclusion in natural language inference
Bill MacCartney and Christopher D. Manning. 2008 · 2008
Earlier work this paper cites.
14. Types of inference: entailment, presupposition, and implicature , pages 397–422. De Gruyter Mouton, Berlin, New York
Yan Huang. 2011 · 2011
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Presupposition: What went wrong?
Lauri Karttunen. 2016 · 2016
Earlier work this paper cites.
The commitmentbank: Investigating projection in naturally occurring discourse
Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019 · 2019
Earlier work this paper cites.
Evaluating BERT for natural language inference: A case study on the CommitmentBank
Nanjiang Jiang and Marie-Catherine de Marneffe. 2019 · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
How well do NLI models capture verb veridicality?
Alexis Ross and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Can neural networks understand monotonicity reasoning?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019a · 2019
Earlier work this paper cites.
HELP: A dataset for identifying shortcomings of neural models in monotonicity reasoning
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019b · 2019
Earlier work this paper cites.
What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models
Allyson Ettinger. 2020 · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Christopher Potts. 2020 · 2020
Earlier work this paper cites.
Probing linguistic systematicity
Emily Goodwin, Koustuv Sinha, and Timothy J. O’Donnell. 2020 · 2020
Earlier work this paper cites.
An analysis of natural language inference benchmarks through the lens of negation
Md Mosharaf Hossain, Venelin Kovatchev, Pranoy Dutta, Tiffany Kao, Elizabeth Wei, and Eduardo Blanco. 2020 · 2020
Earlier work this paper cites.
Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020 · 2020
Cited alongside, same era.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Nora Kassner and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Cited alongside, same era.
Do neural models learn systematicity of monotonicity inference in natural language?
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, and Kentaro Inui. 2020 · 2020
Cited alongside, same era.
A categorical archive of chatgpt failures
Ali Borji. 2023 · 2023
Closest in time.
Language Model Behavior: A Comprehensive Survey
Tyler A. Chang and Benjamin K. Bergen. 2023 · 2023
Closest in time.
It is a bird therefore it is a robin: On BERT’s internal consistency between hypernym knowledge and logical words
Nicolas Guerin and Emmanuel Chemla. 2023 · 2023
Closest in time.
Consistency analysis of chatgpt
Myeongjun Jang and Thomas Lukasiewicz. 2023 · 2023
Closest in time.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Presupposition
David I. Beaver, Bart Geurts, and Kristie Denlinger. 2021 · 2021
Cited alongside, same era.
A multilingual benchmark for probing negation-awareness with minimal pairs
Mareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu, and Anders Søgaard. 2021 · 2021
Cited alongside, same era.
Language models use monotonicity to assess NPI licensing
Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert-Threlkeld. 2021 · 2021
Cited alongside, same era.
NOPE: A corpus of naturally-occurring presuppositions in English
Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, and Tal Linzen. 2021 · 2021
Cited alongside, same era.
Exploring transitivity in neural nli models through veridicality
Hitomi Yanaka, Koji Mineshima, and Kentaro Inui. 2021 · 2021
Cited alongside, same era.
Psycholinguistic diagnosis of language models’ commonsense reasoning
Yan Cong. 2022 · 2022
Cited alongside, same era.
Incremental processing of principle B: Mismatches between neural models and humans
Forrest Davis. 2022 · 2022
Cited alongside, same era.
Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang. 2023 · 2023
Closest in time.
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models
Isabelle Lorge and Janet B. Pierrehumbert. 2023 · 2023
Closest in time.
Vagelis Plevris, George Papazafeiropoulos, and Alejandro Jiménez Rios. 2023 · 2023
Closest in time.
Is chatgpt a general-purpose natural language processing task solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023 · 2023
Closest in time.
In chatgpt we trust? measuring and characterizing the reliability of chatgpt
Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Closest in time.
Language models are not naysayers: an analysis of language models on negation benchmarks
Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023 · 2023
Closest in time.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxin Jiao, Yue Zhang, and Xing Xie. 2023 · 2023
Closest in time.
Assessing step-by-step reasoning against lexical negation: A case study on syllogism
Mengyu Ye, Tatsuki Kuribayashi, Jun Suzuki, Goro Kobayashi, and Hiroaki Funayama. 2023 · 2023
Closest in time.
Can language models be tricked by language illusions? easier with syntax, harder with semantics
Yuhan Zhang, Edward Gibson, and Forrest Davis. 2023 · 2023
Closest in time.
Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023 · 2023
Closest in time.
Dr chatgpt, tell me what i want to hear: How prompt knowledge impacts health answer correctness
Guido Zuccon and Bevan Koopman. 2023 · 2023
Closest in time.
Beyond distributional hypothesis: Let language models learn meaning-text correspondence
Myeongjun Jang, Frank Mtumbuka, and Thomas Lukasiewicz. 2022 · 2042
Closest in time.