Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have become increasingly sophisticated, leading to widespread deployment in sensitive applications where safety and reliability are paramount.
Self-consistency; a theory of personality
Prescott Lecky. 1945 · 1945
Earlier work this paper cites.
Word associations are formed incidentally during sentential semantic integration
Anat Prior and Shlomo Bentin. 2008 · 2008
Earlier work this paper cites.
A simple fine-tuning is all you need: Towards robust deep learning via adversarial fine-tuning
Ahmadreza Jeddi, Mohammad Javad Shafiee, and Alexander Wong. 2020 · 2012
Earlier work this paper cites.
Semantic cosine similarity
Faisal Rahutomo, Teruaki Kitasuka, Masayoshi Aritsugi, et al. 2012 · 2012
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Domain adaptive transfer learning with specialist models
Jiquan Ngiam, Daiyi Peng, Vijay Vasudevan, Simon Kornblith, Quoc V Le, and Ruoming Pang. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Mixture of experts models
Isobel Claire Gormley and Sylvia Frühwirth-Schnatter. 2019 · 2019
Earlier work this paper cites.
Adversarial robustness: From self-supervised pre-training to fine-tuning
Tianlong Chen, Sijia Liu, Shiyu Chang, Yu Cheng, Lisa Amini, and Zhangyang Wang. 2020 · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2020
Earlier work this paper cites.
Self–other moral bias: Evidence from implicit measures and the word-embedding association test
Ming-Hui Li, Pei-Wei Li, and Li-Lin Rao. 2021 · 2021
Earlier work this paper cites.
Decaf: Generating fair synthetic data using causally-aware generative networks
Boris Van Breugel, Trent Kyono, Jeroen Berrevoets, and Mihaela Van der Schaar. 2021 · 2021
Earlier work this paper cites.
Prompt injection: Parameterization of fixed inputs
Eunbi Choi, Yongrae Jo, Joel Jang, and Minjoon Seo. 2022 · 2022
Earlier work this paper cites.
Fine-tuning bert models for intent recognition using a frequency cut-off strategy for domain-specific vocabulary extension
Fernando Fernández-Martínez, Cristina Luna-Jiménez, Ricardo Kleinlein, David Griol, Zoraida Callejas, and Juan Manuel Montero. 2022 · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
A review on fairness in machine learning
Dana Pessach and Erez Shmueli. 2022 · 2022
Earlier work this paper cites.
Intentional biases in llm responses
Nicklaus Badyal, Derek Jacoby, and Yvonne Coady. 2023 · 2023
Cited alongside, same era.
Harms from increasingly agentic algorithmic systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, et al. 2023 · 2023
Cited alongside, same era.
Universal self-consistency for large language model generation
Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Breaking the bias: Gender fairness in llms using prompt engineering and in-context learning
Satyam Dwivedi, Sanjukta Ghosh, and Shivam Dwivedi. 2023 · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023 · 2023
Knowledge graphs as context sources for llm-based explanations of learning recommendations
Hasan Abu-Rasheed, Christian Weber, and Madjid Fathi. 2024 · 2024
Closest in time.
Reliable, adaptable, and attributable language models with retrieval
Akari Asai, Zexuan Zhong, Danqi Chen, Pang Wei Koh, Luke Zettlemoyer, Hannaneh Hajishirzi, and Wen-tau Yih. 2024 · 2024
Closest in time.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, and Sergey Levine. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterexample guided inductive synthesis using large language models and satisfiability solving
Sumit Kumar Jha, Susmit Jha, Patrick Lincoln, Nathaniel D Bastian, Alvaro Velasquez, Rickard Ewetz, and Sandeep Neema. 2023 · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Fairness-guided few-shot prompting for large language models
Huan Ma, Changqing Zhang, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang, Huazhu Fu, Qinghua Hu, and Bingzhe Wu. 2023 · 2023
Cited alongside, same era.
Beyond accuracy: Evaluating self-consistency of code llms
Marcus J Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar, Gail Kaiser, Suman Jana, and Baishakhi Ray. 2023 · 2023
Cited alongside, same era.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 · 2023
Cited alongside, same era.
NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails
Traian Rebedea, Razvan Dinu, Makesh Narsimhan Sreedhar, Christopher Parisien, and Jonathan Cohen. 2023 · 2023
Cited alongside, same era.
Assessing bias in llm-generated synthetic datasets: The case of german voter behavior
Leah von der Heyde, Anna-Carolina Haensch, and Alexander Wenz. 2023 · 2023
Cited alongside, same era.
Xiaoming Liu, Chen Liu, Zhaohan Zhang, Chengzhengxu Li, Longtian Wang, Yu Lan, and Chao Shen. 2024 · 2024
Closest in time.
Towards faithful and robust llm specialists for evidence-based question-answering
Tobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash, and Markus Leippold. 2024 · 2024
Closest in time.
In-context learning agents are asymmetric belief updaters
Johannes A Schubert, Akshay K Jagadish, Marcel Binz, and Eric Schulz. 2024 · 2024
Closest in time.
Fairrag: Fair human generation via fair retrieval augmentation
Robik Shrestha, Yang Zou, Qiuyu Chen, Zhiheng Li, Yusheng Xie, and Siqi Deng. 2024 · 2024
Closest in time.
Systematic biases in llm simulations of debates
Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. 2024 · 2024
Closest in time.
From human experts to machines: An llm supported approach to ontology and knowledge graph construction
Krishna Kommineni Vamsi, Vamsi Krishna Kommineni, and Sheeba Samuel. 2024 · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Closest in time.
Corrective retrieval augmented generation
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024 · 2024
Closest in time.
Matplotagent: Method and evaluation for llm-based agentic scientific data visualization
Zhiyu Yang, Zihan Zhou, Shuo Wang, Xin Cong, Xu Han, Yukun Yan, Zhenghao Liu, Zhixing Tan, Pengyuan Liu, Dong Yu, et al. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024 · 2024
Closest in time.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2024 · 2024
Closest in time.