Fetching the paper…
Reading the bibliography…
Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains.
Counterintuitive behavior of social systems
Jay W Forrester · 1971
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Value sensitive design and responsible innovation
Jeroen Van den Hoven · 2013
Earlier work this paper cites.
Mental models
Dedre Gentner and Albert L Stevens · 2014
Earlier work this paper cites.
Evaluations: autonomy and artificial intelligence: a threat or savior?
William Frere Lawless and Donald A Sofge · 2017
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Causal interpretability for machine learning-problems, methods and evaluation
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Artificial intelligence regulation: a framework for governance
Patricia Gomes Rêgo de Almeida, Carlos Denner dos Santos, and Josivania Silva Farias · 2021
Earlier work this paper cites.
Scaling laws for deep learning
Jonathan S Rosenfeld · 2021
Earlier work this paper cites.
Interpretable counterfactual explanations guided by prototypes
Arnaud Van Looveren and Janis Klaise · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Earlier work this paper cites.
Is power-seeking ai an existential risk?
Joseph Carlsmith · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
https://www.anthropic.com/news/anthropics-responsible-scaling-policy, 2023
Responsible scaling policy · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich · 2023
Earlier work this paper cites.
Statement on ai risk, 2024
Center for AI Safety · 2023
Earlier work this paper cites.
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Earlier work this paper cites.
Flames: Benchmarking value alignment of chinese large language models
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, et al · 2023
Earlier work this paper cites.
Trustworthy ai: From principles to practices
Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou · 2023
Earlier work this paper cites.
Query-relevant images jailbreak large multi-modal models
Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao · 2023
Cited alongside, same era.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI · 2023
Cited alongside, same era.
A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices
Nicola Palladino · 2023
Cited alongside, same era.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Aligning large language models with representation editing: A control perspective
Lingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Yifei Zhou, Yue Song, Rongzhi Zhang, Kai Wang, and Chao Zhang · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao · 2024
Closest in time.
Controllable text generation for large language models: A survey
Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, et al · 2024
Closest in time.
A survey of text watermarking in the era of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural language processing: transforming how machines understand human language (2023)
Abu Rayhan, Robert Kinzler, and Rajan Rayhan · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Cited alongside, same era.
Using the veil of ignorance to align ai systems with principles of justice
Laura Weidinger, Kevin R McKee, Richard Everett, Saffron Huang, Tina O Zhu, Martin J Chadwick, Christopher Summerfield, and Iason Gabriel · 2023
Cited alongside, same era.
Huref: Human-readable fingerprint for large language models
Boyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu, Chenghu Zhou, Xinbing Wang, Yu Yu, and Zhouhan Lin · 2023
Cited alongside, same era.
https://idais.ai/dialogue/idais-beijing/, 2024
Consensus statement on red lines in artificial intelligence · 2024
Cited alongside, same era.
https://idais.ai/dialogue/idais-venice/, 2024
The global nature of ai risks makes it necessary to recognize ai safety as a global public good · 2024
Cited alongside, same era.
Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Xi Zhang, Lijie Wen, Irwin King, Hui Xiong, and Philip Yu · 2024
Closest in time.
Don’t always say no to me: Benchmarking safety-related refusal in large vlm
Xin Liu, Zhichen Dong, Zhanhui Zhou, Yichen Zhu, Yunshi Lan, Jing Shao, Chao Yang, and Yu Qiao · 2024
Closest in time.
Machine unlearning in generative ai: A survey
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang · 2024
Closest in time.
Inference-time language model alignment via integrated value guidance
Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang, Chao Yang, and Yu Qiao · 2024
Closest in time.
Large language models in cybersecurity: State-of-the-art
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd, Pejman Najafi, Feng Cheng, and Christoph Meinel · 2024
Closest in time.
Rule based rewards for language model safety
Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian Kivlichan, Molly Lin, Alex Beutel, John Schulman, and Lilian Weng · 2024
Closest in time.
Accountability in artificial intelligence: what it is and how it works
Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi · 2024
Closest in time.
Sejoon Oh, Yiqiao Jin, Megha Sharma, Donghyun Kim, Eric Ma, Gaurav Verma, and Srijan Kumar · 2024
Closest in time.
Openai o1 system card, 2024
OpenAI · 2024
Closest in time.
Video generation models as world simulators, 2024
OpenAI · 2024
Closest in time.
Automated red teaming with goat: the generative offensive agent tester
Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero, Joanna Bitton, Hailey Nguyen, Joe Li, Cristian Canton Ferrer, Ivan Evtimov, and Aaron Grattafiori · 2024
Closest in time.
Chen Qian, Dongrui Liu, Jie Zhang, Yong Liu, and Jing Shao · 2024
Closest in time.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Identifying semantic induction heads to understand in-context learning
Jie Ren, Qipeng Guo, Hang Yan, Dongrui Liu, Quanshi Zhang, Xipeng Qiu, and Dahua Lin · 2024
Closest in time.
Self-reflection in llm agents: Effects on problem-solving performance. arxiv 2024
M Renze and E Guven · 2024
Closest in time.
Self-reflection in llm agents: Effects on problem-solving performance
Matthew Renze and Erhan Guven · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2024
Closest in time.
Zhen Tao, Dinghao Xi, Zhiyu Li, Liumin Tang, and Wei Xu · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Closest in time.
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al · 2024
Closest in time.
“ai safety as global public goods” working report, 2024
Yingchun Wang, Kai Jia, Jing Zhao, Ling Chen, Chunshen Qin, Yuan Yuan, Hongyu Fu, and Xingzhou Liang · 2024
Closest in time.
Shangyu Xing, Fei Zhao, Zhen Wu, Tuo An, Weihao Chen, Chunhui Li, Jianbing Zhang, and Xinyu Dai · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
The better angels of machine personality: How personality relates to llm safety
Jie Zhang, Dongrui Liu, Chen Qian, Ziyue Gan, Yong Liu, Yu Qiao, and Jing Shao · 2024
Closest in time.
Reef: Representation encoding fingerprints for large language models
Jie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang, Yong Liu, Yu Qiao, and Jing Shao · 2024
Closest in time.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Jing Shao, Xiangyu Yue, Chao Yang, Wanli Ouyang, and Yu Qiao · 2024
Closest in time.
Weak-to-strong search: Align large language models via searching over small language models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao · 2024
Closest in time.
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao · 2025
Closest in time.