Fetching the paper…
Reading the bibliography…
Large language models (LLMs) may exhibit unintended or undesirable behaviors.
The psychology of language
George Kingsley Zipf. 1946 · 1946
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon. 1948 · 1948
Earlier work this paper cites.
A method for the construction of minimum-redundancy codes
David A Huffman. 1952 · 1952
Earlier work this paper cites.
Cooperation under the security dilemma
Robert Jervis. 1978 · 1978
Earlier work this paper cites.
Arithmetic coding for data compression
Ian H Witten, Radford M Neal, and John G Cleary. 1987 · 1987
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Universal Artificial Intelligence: Sequential Decisions Based On Algorithmic Probability
Marcus Hutter. 2005 · 2005
Earlier work this paper cites.
Power laws, pareto distributions and zipf’s law
Mark EJ Newman. 2005 · 2005
Earlier work this paper cites.
Is this paper dangerous? balancing secrecy and openness in counterterrorism
Jacob N Shapiro and David A Siegel. 2010 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Lectures de potentia restitutiva, or of spring explaining the power of springing bodies
Robert Hooke. 2016 · 2016
Earlier work this paper cites.
A survey of transfer learning
Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang. 2016 · 2016
Earlier work this paper cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Earlier work this paper cites.
How does the offense-defense balance scale?
Ben Garfinkel and Allan Dafoe. 2021 · 2021
Earlier work this paper cites.
Artificial intelligence act
Tambiama Madiega. 2021 · 2021
Earlier work this paper cites.
Is sgd a bayesian sampler? well, almost
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse, and Ard A Louis. 2021 · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Dual use of artificial-intelligence-powered drug discovery
Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphael Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. 2023 · 2023
Earlier work this paper cites.
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. 2023 · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, SHUM KaShun, and Tong Zhang. 2023 · 2023
Earlier work this paper cites.
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova. 2023 · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023 · 2023
Cited alongside, same era.
More than a feeling: Accuracy and application of sentiment analysis
Jochen Hartmann, Mark Heitmann, Christian Siebert, and Christina Schamp. 2023 · 2023
Cited alongside, same era.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger. 2023 · 2023
Cited alongside, same era.
Ai alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. 2023 · 2023
Cited alongside, same era.
Consensus statement on red lines in artificial intelligence
IDAIS-Beijing. 2024 · 2024
Closest in time.
A literature survey on open source large language models
Sanjay Kukreja, Tarun Kumar, Amit Purohit, Abhijit Dasgupta, and Debashis Guha. 2024 · 2024
Closest in time.
When your ais deceive you: Challenges of partial observability in reinforcement learning from human feedback
Leon Lang, Davis Foote, Stuart J Russell, Anca Dragan, Erik Jenner, and Scott Emmons. 2024 · 2024
Closest in time.
Rrm: Robust reward model training mitigates reward hacking
Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, et al. 2024 · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Cited alongside, same era.
Remax: A simple, effective, and efficient method for aligning large language models
Ziniu Li, Tian Xu, Yushun Zhang, Yang Yu, Ruoyu Sun, and Zhi-Quan Luo. 2023 · 2023
Cited alongside, same era.
Probabilistic machine learning: Advanced topics
Kevin P Murphy. 2023 · 2023
Cited alongside, same era.
Jonas B Sandbrink. 2023 · 2023
Cited alongside, same era.
Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, et al. 2023 · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Cited alongside, same era.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Jing Shao, Xiangyu Yue, Chao Yang, Wanli Ouyang, and Yu Qiao. 2023 · 2023
Cited alongside, same era.
Auditing large language models: a three-layered approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2024 · 2024
Closest in time.
Ai deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2024 · 2024
Closest in time.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Open problems in technical ai governance
Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, et al. 2024 · 2024
Closest in time.
Seal: Systematic error analysis for value alignment
Manon Revel, Matteo Cargnelutti, Tyna Eloundou, and Greg Leppert. 2024 · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. 2024 · 2024
Closest in time.
Language models learn to mislead humans via rlhf
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R Bowman, He He, and Shi Feng. 2024 · 2024
Closest in time.
Chaojun Xiao, Jie Cai, Weilin Zhao, Guoyang Zeng, Biyuan Lin, Jie Zhou, Zhi Zheng, Xu Han, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. 2024 · 2024
Closest in time.
Alignment for honesty
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2024 · 2024
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024 · 2024
Closest in time.
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation
Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi. 2025 · 2025
Closest in time.
International ai safety report
Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. 2025 · 2025
Closest in time.
Mixpro: Simple yet effective data augmentation for prompt-based learning
Bohan Li, Longxu Dou, Yutai Hou, Yunlong Feng, Honglin Mu, Enbo Wang, Qingfu Zhu, Qinghua Sun, and Wanxiang Che. 2025 · 2025
Closest in time.
Against the achilles’ heel: A survey on red teaming for generative models
Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang, Renxi Wang, Junjie Gao, Yixuan Zhang, Wanxiang Che, Timothy Baldwin, et al. 2025 · 2025
Closest in time.
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra-Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, et al. 2025 · 2025
Closest in time.
A survey of table reasoning with large language models
Xuanliang Zhang, Dingzirui Wang, Longxu Dou, Qingfu Zhu, and Wanxiang Che. 2025 · 2025
Closest in time.
Sequence to sequence reward modeling: Improving rlhf by language feedback
Jiayi Zhou, Jiaming Ji, Josef Dai, and Yaodong Yang. 2025 · 2025
Closest in time.