Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly explored as knowledge bases (KBs), yet current evaluation methods focus too narrowly on knowledge retention, overlooking other crucial criteria for reliable performance.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Self-alignment for factuality: Mitigating hallucinations in LLMs via self-evaluation
Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, and Helen Meng. 2024b · 1965
Earlier work this paper cites.
Easy cases of probabilistic satisfiability
Kim Allan Andersen and Daniele Pretolani. 2001 · 2001
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Trustscore: Reference-free evaluation of llm response trustworthiness
Danna Zheng, Danyang Liu, Mirella Lapata, and Jeff Z Pan. 2024 · 2019
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text models
Rakesh Chada and Pradeep Natarajan. 2021 · 2021
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
Can generative pre-trained language models serve as knowledge bases for closed-book qa?
Cunxiang Wang, Pai Liu, and Yue Zhang. 2021 · 2021
Earlier work this paper cites.
Becel: Benchmark for consistency evaluation of language models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz. 2022 · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Cited alongside, same era.
Felm: Benchmarking factuality evaluation of large language models
Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He. 2023 · 2023
Cited alongside, same era.
The effect of scaling, retrieval augmentation and form on the factual consistency of language models
Lovisa Hagström, Denitsa Saynova, Tobias Norlund, Moa Johansson, and Richard Johansson. 2023 · 2023
Can language models act as knowledge bases at scale?
Qiyuan He, Yizhong Wang, and Wenya Wang. 2024 · 2024
Closest in time.
Evaluating the factuality of large language models using large-scale knowledge graphs
Xiaoze Liu, Feijie Wu, Tianyang Xu, Zhuo Chen, Yichi Zhang, Xiaoqian Wang, and Jing Gao. 2024 · 2024
Closest in time.
Generating benchmarks for factuality evaluation of language models
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham. 2024 · 2024
Closest in time.
Why does new knowledge create messy ripple effects in llms?
Jiaxin Qin, Zixuan Zhang, Chi Han, Manling Li, Pengfei Yu, and Heng Ji. 2024 · 2024
Closest in time.
Knowledge-based consistency testing of large language models
Sai Sathiesh Rajan, Ezekiel Soremekun, and Sudipta Chattopadhyay. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al. 2023 · 2023
Cited alongside, same era.
Kai Sun, Yifan Ethan Xu, Hanwen Zha, Yue Liu, and Xin Luna Dong. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024 · 2024
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2024 · 2024
Cited alongside, same era.
Detoxifying large language models via knowledge editing
Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. 2024b
Cited in the paper.
Closest in time.
Evaluating consistency and reasoning capabilities of large language models
Yash Saxena, Sarthak Chopra, and Arunendra Mani Tripathi. 2024 · 2024
Closest in time.
AXCEL: Automated eXplainable consistency evaluation using LLMs
P Aditya Sreekar, Sahil Verma, Suransh Chopra, Abhishek Persad, Sarik Ghazarian, and Narayanan Sadagopan. 2024 · 2024
Closest in time.
RoseLoRA: Row and column-wise sparse low-rank adaptation of pre-trained language model for knowledge editing and fine-tuning
Haoyu Wang, Tianci Liu, Ruirui Li, Monica Xiao Cheng, Tuo Zhao, and Jing Gao. 2024a · 2024
Closest in time.
Factuality of large language models in the year 2024
Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Georgi Georgiev, Rocktim Jyoti Das, and Preslav Nakov. 2024d · 2024
Closest in time.
Truth-aware context selection: Mitigating hallucinations of large language models being misled by untruthful contexts
Tian Yu, Shaolei Zhang, and Yang Feng. 2024 · 2024
Closest in time.
Felm: Benchmarking factuality evaluation of large language models
Yiran Zhao, Jinghan Zhang, I Chern, Siyang Gao, Pengfei Liu, Junxian He, et al. 2024 · 2024
Closest in time.