Fetching the paper…
Reading the bibliography…
LLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Reclor: A reading comprehension dataset requiring logical reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Transcending scaling laws with 0.1% extra compute
Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q Tran, David R So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, et al · 2022
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Earlier work this paper cites.
Controlled text generation via language model arithmetic
Jasper Dekoninck, Marc Fischer, Luca Beurer-Kellner, and Martin Vechev · 2023
Earlier work this paper cites.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Earlier work this paper cites.
Can large language models truly understand prompts? a case study with negated prompts
Joel Jang, Seonghyeon Ye, and Minjoon Seo · 2023
Earlier work this paper cites.
Improving activation steering in language models with mean-centring
Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan · 2023
Earlier work this paper cites.
Specific versus general principles for constitutional ai
Sandipan Kundu, Yuntao Bai, Saurav Kadavath, Amanda Askell, Andrew Callahan, Anna Chen, Anna Goldie, Avital Balwit, Azalia Mirhoseini, Brayden McLean, et al · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Earlier work this paper cites.
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
X Liu, Y Zhu, J Gu, Y Lan, C Yang, and Y Qiao · 2023
Earlier work this paper cites.
Socialstigmaqa: A benchmark to uncover stigma amplification in generative language models, 2023
Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini · 2023
Earlier work this paper cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Earlier work this paper cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Earlier work this paper cites.
Failure modes of learning reward models for llms and other sequence models
Silviu Pitis · 2023
Earlier work this paper cites.
I’m afraid i can’t do that: Predicting prompt refusal in black-box generative language models
Max Reuter and William Schulze · 2023
Earlier work this paper cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Earlier work this paper cites.
Arb: Advanced reasoning benchmark for large language models, 2023
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, and Aran Komatsuzaki · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Zephyr: Direct distillation of lm alignment, 2023
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Cited alongside, same era.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric. P Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Cited alongside, same era.
Investigating bias representations in llama 2 chat via activation steering, 2024
Dawn Lu and Nina Rimsky · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta · 2024
Closest in time.
Sample-efficient preference-based reinforcement learning with dynamics aware rewards
Katherine Metcalf, Miguel Sarabia, Natalie Mackraz, and Barry-John Theobald · 2024
Closest in time.
Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin, Gabor Fodor, Nischay Dhankhar, and Sri Satish Ambati · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dyah Adila, Shuai Zhang, Boran Han, and Yuyang Wang · 2024
Cited alongside, same era.
Scaling sparse fine-tuning to large language models
Alan Ansell, Ivan Vulić, Hannah Sterz, Anna Korhonen, and Edoardo M Ponti · 2024
Cited alongside, same era.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Understanding jailbreak success: A study of latent space dynamics in large language models
Sarah Ball, Frauke Kreuter, and Nina Rimsky · 2024
Cited alongside, same era.
The art of saying no: Contextual noncompliance in language models
Faeze Brahman, Sachin Kumar, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Chandu, Jack Hessel, et al · 2024
Cited alongside, same era.
Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin, Lu Lin, Fenglong Ma, and Jinghui Chen · 2024
Cited alongside, same era.
Ziwei Chai, Guoyin Wang, Jing Su, Tianjie Zhang, Xuanwen Huang, Xuwu Wang, Jingjing Xu, Jianbo Yuan, Hongxia Yang, Fei Wu, et al · 2024
Cited alongside, same era.
Closest in time.
Phuc Phan, Hieu Tran, and Long Phan · 2024
Closest in time.
Spectral editing of activations for large language model alignment
Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo M Ponti, and Shay B Cohen · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Controlling large language model agents with entropic activation steering
Nate Rahn, Pierluca D’Oro, and Marc G Bellemare · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
Multi-property steering of large language models with dynamic activation composition
Daniel Scalena, Gabriele Sarti, and Malvina Nissim · 2024
Closest in time.
Transformers represent belief state geometry in their residual stream
Adam S Shai, Sarah E Marzen, Lucas Teixeira, Alexander Gietelink Oldenziel, and Paul M Riechers · 2024
Closest in time.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Closest in time.
Ondrej Sotolar · 2024
Closest in time.
Steering without side effects: Improving post-deployment control of language models
Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau, Salsabila Mahdi, and Samuel R Bowman · 2024
Closest in time.
Lab: Large-scale alignment for chatbots
Shivchander Sudalairaj, Abhishek Bhandwaldar, Aldo Pareja, Kai Xu, David D Cox, and Akash Srivastava · 2024
Closest in time.
Llm roleplay: Simulating human-chatbot interaction
Hovhannes Tamoyan, Hendrik Schuff, and Iryna Gurevych · 2024
Closest in time.
Analyzing the generalization and reliability of steering vectors
Daniel Chee Hian Tan, David Chanin, Aengus Lynch, Adrià Garriga-Alonso, Dimitrios Kanoulas, Brooks Paige, and Robert Kirk · 2024
Closest in time.
Hermes-2-pro-llama-3-8b, 2024
Teknium, interstellarninja, theemozilla, karan4d, and huemin_art · 2024
Closest in time.
Exploring and steering the moral compass of large language models
Alejandro Tlaie · 2024
Closest in time.
Does editing provide evidence for localization?
Zihao Wang and Victor Veitch · 2024
Closest in time.
The art of refusal: A survey of abstention in large language models
Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang · 2024
Closest in time.
Reft: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts · 2024
Closest in time.
Wizardlm: Empowering large pre-trained language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang · 2024
Closest in time.
Lofit: Localized fine-tuning on llm representations
Fangcong Yin, Xi Ye, and Greg Durrett · 2024
Closest in time.
The better angels of machine personality: How personality relates to llm safety
Jie Zhang, Dongrui Liu, Chen Qian, Ziyue Gan, Yong Liu, Yu Qiao, and Jing Shao · 2024
Closest in time.
Prompt-driven llm safeguarding via directed representation optimization
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.