Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved impressive advancements across numerous disciplines, yet the critical issue of knowledge conflicts, a major source of hallucinations, has rarely been studied.
Wikidata: A free collaborative knowledge base
Denny Vrandečić and Markus Krötzsch · 2014
Earlier work this paper cites.
Entity-based knowledge conflicts in question answering
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh · 2021
Earlier work this paper cites.
Wikicontradiction: Detecting self-contradiction articles on wikipedia
Cheng Hsu, Cheng-Te Li, Diego Saez-Trumper, and Yi-Zhan Hsu · 2021
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, et al · 2021
Earlier work this paper cites.
PolyLM: Learning about polysemy through language modeling
Alan Ansell, Felipe Bravo-Marquez, and Bernhard Pfahringer · 2021
Earlier work this paper cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Earlier work this paper cites.
A dataset for answering time-sensitive questions, 2021
Wenhu Chen, Xinyi Wang, and William Yang Wang · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Situatedqa: Incorporating extra-linguistic contexts into qa
Michael JQ Zhang and Eunsol Choi · 2021
Earlier work this paper cites.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg · 2021
Earlier work this paper cites.
{ \{ ZeRO-Offload } \} : Democratizing { \{ Billion-Scale } \} model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le · 2022
Earlier work this paper cites.
Claimdiff: Comparing and contrasting claims on contentious issues
Miyoung Ko, Ingyu Seong, Hwaran Lee, Joonsuk Park, Minsuk Chang, and Minjoon Seo · 2022
Earlier work this paper cites.
Synthetic disinformation attacks on automated fact verification systems
Yibing Du, Antoine Bosselut, and Christopher D Manning · 2022
Earlier work this paper cites.
Improving temporal generalization of pre-trained language models with lexical semantic change
Zhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu, Min Zhang, and Juntao Li · 2022
Earlier work this paper cites.
Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence
Hung-Ting Chen, Michael Zhang, and Eunsol Choi · 2022
Earlier work this paper cites.
A survey on evaluation of large language models, 2023
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie · 2023
Earlier work this paper cites.
On the risk of misinformation pollution with large language models, 2023
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang · 2023
Earlier work this paper cites.
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su · 2023
Cited alongside, same era.
Context-faithful prompting for large language models
Wenxuan Zhou, Sheng Zhang, Hoifung Poon, and Muhao Chen · 2023
Cited alongside, same era.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Synthetic prompting: Generating chain-of-thought demonstrations for large language models
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen · 2023
Cited alongside, same era.
Alignment for honesty, 2023
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Later among the works it cites.
Efficient continue training of temporal language model with structural information
Zhaochen Su, Juntao Li, Zikang Zhang, Zihan Zhou, and Min Zhang · 2023
Later among the works it cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Later among the works it cites.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du · 2024
Closest in time.
Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Contradoc: understanding self-contradictions in documents with large language models
Jierui Li, Vipul Raheja, and Dhruv Kumar · 2023
Cited alongside, same era.
Check your facts and try again: Improving large language models with external knowledge and automated feedback, 2023
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao · 2023
Cited alongside, same era.
Language model behavior: A comprehensive survey, 2023
Tyler A. Chang and Benjamin K. Bergen · 2023
Cited alongside, same era.
Predicting question-answering performance of large language models through semantic consistency, 2023
Ella Rabinovich, Samuel Ackerman, Orna Raz, Eitan Farchi, and Ateret Anaby-Tavor · 2023
Cited alongside, same era.
Self-consistency of large language models under ambiguity, 2023
Henning Bartsch, Ole Jorgensen, Domenic Rosati, Jason Hoelscher-Obermaier, and Jacob Pfau · 2023
Cited alongside, same era.
Combating misinformation in the age of llms: Opportunities and challenges, 2023
Canyu Chen and Kai Shu · 2023
Cited alongside, same era.
What evidence do language models find convincing?
Alexander Wan, Eric Wallace, and Dan Klein · 2024
Closest in time.
Twin-merging: Dynamic integration of modular expertise in model merging
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, and Yu Cheng · 2024
Closest in time.
Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang, Qi Cao, and Xueqi Cheng · 2024
Closest in time.
Mathattack: Attacking large language models towards math solving ability
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang · 2024
Closest in time.
Confidence is not timeless: Modeling temporal validity for rule-based temporal knowledge graph forecasting
Rikui Huang, Wei Wei, Xiaoye Qu, Shengzhe Zhang, Dangyang Chen, and Yu Cheng · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?, 2024
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig · 2024
Closest in time.
Realtime qa: What’s the answer right now?
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, Kentaro Inui, et al · 2024
Closest in time.
Managing extreme ai risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, and Sören Mindermann · 2024
Closest in time.
Genai against humanity: nefarious applications of generative artificial intelligence and large language models
Emilio Ferrara · 2024
Closest in time.
Intuitive or dependent? investigating llms’ behavior style to conflicting prompts, 2024
Jiahao Ying, Yixin Cao, Kai Xiong, Yidong He, Long Cui, and Yongbin Liu · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, and Yongqiang Ma · 2024
Closest in time.
Llama-moe: Building mixture-of-experts from llama with continual pre-training
Tong Zhu, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, and Yu Cheng · 2024
Closest in time.