Fetching the paper…
Reading the bibliography…
While large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs".
Using weighted MAX-SAT engines to solve MPE
James D Park. 2002 · 2002
Earlier work this paper cites.
An introduction to factor graphs
Hans-Andrea Loeliger. 2004 · 2004
Earlier work this paper cites.
A dynamic approach for MPE and weighted MAX-SAT
Tian Sang, Paul Beame, and Henry A Kautz. 2007 · 2007
Earlier work this paper cites.
Transforming question answering datasets into natural language inference datasets
Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander Rush. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
RC2: an efficient MaxSAT solver
Alexey Ignatiev. 2019 · 2019
Earlier work this paper cites.
A logic-driven framework for consistency of neural models
Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar. 2019 · 2019
Earlier work this paper cites.
Knowledge enhanced contextual word representations
Matthew E Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A Smith. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
QuaRTz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019 · 2019
Earlier work this paper cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Cited alongside, same era.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Nora Kassner and H. Schütze. 2020 · 2020
Cited alongside, same era.
How context affects language models’ factual predictions
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020 · 2020
Cited alongside, same era.
On the systematicity of probing contextualized word representations: The case of hypernymy in BERT
Abhilasha Ravichander, Eduard Hovy, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung. 2020 · 2020
Cited alongside, same era.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, T. J. Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zachary Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yushi Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, John Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom B. Brown, Jack Clark, Nicholas Joseph, Benjamin Mann, Sam McCandlish, Christopher Olah, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Towards faithful model explanation in NLP: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2022 · 2022
Later among the works it cites.
Enhancing self-consistency and performance of pre-trained language models through natural language inference
Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher D. Manning. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Cited alongside, same era.
Obtaining faithful interpretations from compositional neural networks
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh, Jonathan Berant, and Matt Gardner. 2020 · 2020
Cited alongside, same era.
Explaining answers with entailment trees
Bhavana Dalvi, Peter Alexander Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021 · 2021
Cited alongside, same era.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, E. Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Nora Kassner, Oyvind Tafjord, Hinrich Schutze, and Peter Clark. 2021 · 2021
Cited alongside, same era.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasović. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
ChatGPT: Optimizing language models for dialog
OpenAI. 2022 · 2022
Later among the works it cites.
Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs
Kyle Richardson, Ronen Tamari, Oren Sultan, Reut Tsarfaty, Dafna Shahaf, and Ashish Sabharwal. 2022 · 2022
Later among the works it cites.
Entailer: Answering questions with faithful and truthful chains of reasoning
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Dynamic generation of interpretable inference rules in a neuro-symbolic expert system
Nathaniel Weir and Benjamin Van Durme. 2022 · 2022
Later among the works it cites.
Do language models have coherent mental models of everyday things?
Yuling Gu, Bhavana Dalvi, and Peter Clark. 2023 · 2023
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Scharli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Huai hsin Chi. 2023 · 2023
Closest in time.