Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) often exhibit knowledge disparities across languages.
Models, reasoning and inference
Judea Pearl et al. 2000 · 2000
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Earlier work this paper cites.
Causalbert: Injecting causal knowledge into pre-trained models with minimal supervision
Zhongyang Li, Xiao Ding, Kuo Liao, Bing Qin, and Ting Liu. 2021 · 2021
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021 · 2021
Earlier work this paper cites.
Cross-cultural similarity features for cross-lingual transfer learning of pragmatically motivated tasks
Jimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, and David R. Mortensen. 2021 · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
The internal state of an LLM knows when it‘s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Earlier work this paper cites.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023 · 2023
Earlier work this paper cites.
Faithful explanations of black-box nlp models using llm-generated counterfactuals
Yair Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Amit Sharma, and Roi Reichart. 2023 · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Cited alongside, same era.
Language generation models can cause harm: So what can we do about it? an actionable survey
Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
Gpt-4 technical report. arxiv 2303.08774
R OpenAI. 2023 · 2023
Cited alongside, same era.
Getting more out of mixture of language model reasoning experts
Chenglei Si, Weijia Shi, Chen Zhao, Luke Zettlemoyer, and Jordan Boyd-Graber. 2023 · 2023
Cited alongside, same era.
The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023 · 2023
Black-box prompt optimization: Aligning large language models without model training
Jiale Cheng, Xiao Liu, Kehan Zheng, Pei Ke, Hongning Wang, Yuxiao Dong, Jie Tang, and Minlie Huang. 2024 · 2024
Later among the works it cites.
Teaching LLMs to abstain across languages via multilingual feedback
Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, and Yulia Tsvetkov. 2024b · 2024
Later among the works it cites.
Llm4causal: Democratized causal tools for everyone via large language model
Haitao Jiang, Lin Ge, Yuhe Gao, Jianian Wang, and Rui Song. 2024 · 2024
Later among the works it cites.
Comparing hallucination detection metrics for multilingual generation
Haoqiang Kang, Terra Blevins, and Luke Zettlemoyer. 2024 · 2024
Later among the works it cites.
Do llms know when to not answer? investigating abstention abilities of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 · 2023
Cited alongside, same era.
Post-abstention: Towards reliably re-attempting the abstained instances in qa
Neeraj Varshney and Chitta Baral. 2023 · 2023
Cited alongside, same era.
Kola: Carefully benchmarking world knowledge of large language models
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, et al. 2023 · 2023
Cited alongside, same era.
Discovering the real association: Multimodal causal reasoning in video question answering
Chuanqi Zang, Hanqing Wang, Mingtao Pei, and Wei Liang. 2023 · 2023
Cited alongside, same era.
Don‘t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs
Xiang Zhang, Senyu Li, Bradley Hauer, Ning Shi, and Grzegorz Kondrak. 2023a · 2023
Cited alongside, same era.
Causal-cog: A causal-effect look at context generation for boosting multi-modal language models
Shitian Zhao, Zhuowan Li, Yadong Lu, Alan Yuille, and Yan Wang. 2023 · 2023
Cited alongside, same era.
Learning a structural causal model for intuition reasoning in conversation
Hang Chen, Bingyu Liao, Jing Luo, Wenjing Zhu, and Xinyu Yang. 2024 · 2024
Cited alongside, same era.
Nishanth Madhusudhan, Sathwik Tejaswi Madhusudhan, Vikas Yadav, and Masoud Hashemi. 2024 · 2024
Later among the works it cites.
Fine-grained hallucination detection and editing for language models
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi. 2024 · 2024
Later among the works it cites.
A survey of large language models on generative graph analytics: Query, learning, and applications
Wenbo Shang and Xin Huang. 2024 · 2024
Later among the works it cites.
DeCoT: Debiasing chain-of-thought for knowledge-intensive tasks in large language models via causal intervention
Junda Wu, Tong Yu, Xiang Chen, Haoliang Wang, Ryan Rossi, Sungchul Kim, Anup Rao, and Julian McAuley. 2024 · 2024
Later among the works it cites.
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024 · 2024
Later among the works it cites.
The knowledge alignment problem: Bridging human and external knowledge for large language models
Shuo Zhang, Liangming Pan, Junzhou Zhao, and William Yang Wang. 2024 · 2024
Later among the works it cites.
Batch calibration: Rethinking calibration for in-context learning and prompt engineering
Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine Heller, and Subhrajit Roy. 2024 · 2024
Later among the works it cites.
Multimodal commonsense knowledge distillation for visual question answering (student abstract)
Shuo Yang, Siwen Luo, and Soyeon Caren Han. 2025 · 2025
Closest in time.