Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consistency.
Learning to parse database queries using inductive logic programming
John M. Zelle and Raymond J. Mooney · 1996
Earlier work this paper cites.
Learning to transform natural to formal languages
Rohit J. Kate, Yuk Wah Wong, and Raymond J. Mooney · 2005
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Measuring and Improving Consistency in Pretrained Language Models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
Accurate, yet inconsistent? consistency analysis on language understanding models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz · 2021
Earlier work this paper cites.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?, 2021
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2021
Cited alongside, same era.
Factual consistency of multilingual pretrained language models
Constanza Fierro and Anders Søgaard · 2022
Cited alongside, same era.
BECEL: Benchmark for consistency evaluation of language models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz · 2022
Cited alongside, same era.
Enhancing self-consistency and performance of pretrained language models with nli
Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher D. Manning · 2022
Cited alongside, same era.
Measuring reliability of large language models through semantic consistency
Harsh Raj, Domenic Rosati, and Subhabrata Majumdar · 2022
Later among the works it cites.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2022
Later among the works it cites.
Prompt consistency for zero-shot task generalization
Chunting Zhou, Junxian He, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig · 2022
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality, 2023
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Newman, Prafulla Kumar Choubey, and Nazneen Rajani · 2022
Cited alongside, same era.
Do the openai api models have knowledge of current events?
OpenAI Help Center
Cited in the paper.
Model index for researchers
OpenAI
Cited in the paper.
Wikimedia downloads
Wikimedia Foundation
Cited in the paper.
Large language models sensitivity to the order of options in multiple-choice questions, 2023
Pouya Pezeshkpour and Estevam Hruschka · 2023
Closest in time.