Fetching the paper…
Reading the bibliography…
Large language models (LLMs) pose significant risks due to the potential for generating harmful content or users attempting to evade guardrails.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Posterior calibration and exploratory analysis for natural language processing models
Khanh Nguyen and Brendan O’Connor · 2015
Earlier work this paper cites.
Towards bayesian deep learning: A framework and some existing methods
Hao Wang and Dit-Yan Yeung · 2016
Earlier work this paper cites.
Natural-parameter networks: A class of probabilistic neural networks
Hao Wang, SHI Xingjian, and Dit-Yan Yeung · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
A survey on bayesian deep learning
Hao Wang and Dit-Yan Yeung · 2020
Earlier work this paper cites.
Continuously indexed domain adaptation
Hao Wang, Hao He, and Dina Katabi · 2020
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Earlier work this paper cites.
Revisiting the calibration of modern neural networks
Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
A close look into the calibration of pre-trained language models
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji · 2022
Earlier work this paper cites.
Extrapolative continuous-time bayesian neural network for fast training-free test-time adaptation
Hengguan Huang, Xiangming Gu, Hao Wang, Chang Xiao, Hongfu Liu, and Ye Wang · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers
Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman · 2022
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Graph-relational domain adaptation
Zihao Xu, Guang-He Lee, Yuyang Wang, Hao Wang, et al · 2022
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang · 2023
Cited alongside, same era.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Later among the works it cites.
On the calibration of large language models and alignment
Chiwei Zhu, Benfeng Xu, Quan Wang, Yongdong Zhang, and Zhendong Mao · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Enhancing in-context learning via linear probe calibration
Momin Abbas, Yi Zhou, Parikshit Ram, Nathalie Baracaldo, Horst Samulowitz, Theodoros Salonidis, and Tianyi Chen · 2024
Closest in time.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
Mitigating label biases for in-context learning
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Cited alongside, same era.
Towards informative few-shot prompt with maximum information gain for in-context learning
Hongfu Liu and Ye Wang · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Cited alongside, same era.
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien · 2024
Closest in time.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Blob: Bayesian low-rank adaptation by backpropagation for large language models
Yibin Wang, Haizhou Shi, Ligong Han, Dimitris Metaxas, and Hao Wang · 2024
Closest in time.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content
Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Song, and Bo Li · 2024
Closest in time.
Shieldgemma: Generative ai content moderation based on gemma
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, et al · 2024
Closest in time.