Fetching the paper…
Reading the bibliography…
Language models (LMs) have been shown to behave unexpectedly post-deployment.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models, October 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan · 2022
Earlier work this paper cites.
Locating and Editing Factual Associations in GPT, June 2022
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Extracting Latent Steering Vectors from Pretrained Language Models, May 2022
Nishant Subramani, Nivedita Suresh, and Matthew E. Peters · 2022
Earlier work this paper cites.
Frontier AI Regulation: Managing Emerging Risks to Public Safety, 2023
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf · 2023
Earlier work this paper cites.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations, May 2023
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou · 2023
Earlier work this paper cites.
Inspecting and Editing Knowledge Representations in Language Models, May 2023
Evan Hernandez, Belinda Z. Li, and Jacob Andreas · 2023
Earlier work this paper cites.
Improving Activation Steering in Language Models with Mean-Centring
Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan · 2023
Cited alongside, same era.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
A Conversation With Bing’s Chatbot Left Me Deeply Unsettled, 2023
Kevin Roose · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, and Kevin Stone · 2023
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought, 2024
James Chua, Edward Rees, Hunar Batra, Samuel R. Bowman, Julian Michael, Ethan Perez, and Miles Turpin · 2024
Closest in time.
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
A Trivial Jailbreak Against Llama 3
Haize Labs · 2024
Closest in time.
Investigating Bias Representations in Llama 2 Chat via Activation Steering, 2024
Dawn Lu and Nina Rimsky · 2024
Closest in time.
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Cited alongside, same era.
Haoran Wang and Kai Shu · 2023
Cited alongside, same era.
Jailbroken: How Does LLM Safety Training Fail?, July 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, 2024
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Refusal in Language Models Is Mediated by a Single Direction, 2024
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Mixtral of Experts, January 2024a
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed
Cited in the paper.
Closest in time.
Steering Llama 2 via Contrastive Activation Addition, 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
Meta Llama Guard 2
Llama Team · 2024
Closest in time.
A Language Model’s Guide Through Latent Space, 2024
Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2024
Closest in time.
SWE-agent: Agent Computer Interfaces Enable Software Engineering Language Models, 2024
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Closest in time.
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report, 2024
Justin Zhao, Timothy Wang, Wael Abid, Geoffrey Angus, Arnav Garg, Jeffery Kinnison, Alex Sherstinsky, Piero Molino, Travis Addair, and Devvret Rishi · 2024
Closest in time.