Fetching the paper…
Reading the bibliography…
Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes.
The curious case of neural text degeneration, 2020, arXiv preprint:1904.09751
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 1904
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models, 2020, arXiv preprint:2009.11462
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2009
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification, 2015, arXiv preprint:1502.01852
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Information dropout: Learning optimal representations through noisy computation
Alessandro Achille and Stefano Soatto · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Dynamics and reachability of learning tasks, 2019, arXiv preprint:1810.02440
Alessandro Achille, Glen Mbeng, and Stefano Soatto · 2019
Earlier work this paper cites.
Constrained decoding for neural NLG from compositional representations in task-oriented dialogue
Anusha Balakrishnan, Jinfeng Rao, Kartikeya Upasani, Michael White, and Rajen Subba · 2019
Earlier work this paper cites.
Semantic Noise Matters for Neural Natural Language Generation
Ondřej Dušek, David M Howcroft, and Verena Rieser · 2019
Earlier work this paper cites.
ViGGO: A video game corpus for data-to-text generation in open-domain conversation
Juraj Juraska, Kevin Bowden, and Marilyn Walker · 2019
Earlier work this paper cites.
Learnability for the information bottleneck
Tailin Wu, Ian Fischer, Isaac L. Chuang, and Max Tegmark · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
CrowS-Pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman · 2020
Earlier work this paper cites.
The CACAPO dataset: A multilingual, multi-domain dataset for neural pipeline and end-to-end data-to-text generation
Chris van der Lee, Chris Emmery, Sander Wubben, and Emiel Krahmer · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021, arXiv preprint:2106.09685
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Unlearnable examples: Making personal data unexploitable, 2021, arXiv preprint:2101.04898
Hanxun Huang, Xingjun Ma, Sarah Monazam Erfani, James Bailey, and Yisen Wang · 2021
Earlier work this paper cites.
DART: Open-domain structured data record to text generation
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani · 2021
Earlier work this paper cites.
Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization
Lixu Wang, Shichao Xu, Ruiqi Xu, Xiao Wang, and Qi Zhu · 2021
Earlier work this paper cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Yonatan Belinkov · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022, arXiv preprint:2209.07858
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers, 2022, arXiv preprint:2202.11176
Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
Qwen technical report, 2023, arXiv preprint:2309.16609
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023, arXiv preprint:2303.08112
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Language model unalignment: Parametric red-teaming to expose hidden harms and biases
Rishabh Bhardwaj and Soujanya Poria · 2023
Universal and transferable adversarial attacks on aligned language models, 2023, arXiv preprint:2307.15043
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Closest in time.
Stealing part of a production language model, 2024, arXiv preprint:2403.06634
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, and Florian Tramèr · 2024
Closest in time.
Defending against unforeseen failure modes with latent adversarial training, 2024, arXiv preprint:2403.05030
Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell · 2024
Closest in time.
SOPHON: Non-fine-tunable learning to restrain task transferability for pre-trained models, 2024, arXiv preprint:2404.12699
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Purple llama cyberseceval: A secure coding benchmark for language models, 2023, arXiv preprint:2312.04724
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe · 2023
Cited alongside, same era.
Hazards from increasingly accessible fine-tuning of downloadable foundation models, 2023, arXiv preprint:2312.14751
Alan Chan, Ben Bucknall, Herbie Bradley, and David Krueger · 2023
Cited alongside, same era.
Robbie: Robust bias evaluation of large generative language models, 2023, arXiv preprint:2311.18140
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Michael Smith · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Cited alongside, same era.
Will releasing the weights of future large language models grant widespread access to pandemic agents?, 2023, arXiv preprint:2310.18233
Anjali Gopal, Nathan Helm-Burger, Lennart Justen, Emily H. Soice, Tiffany Tzeng, Geetha Jeyapragasan, Simon Grimm, Benjamin Mueller, and Kevin M. Esvelt · 2023
Cited alongside, same era.
DeBERTaV3: Improving DeBERTa using ELECTRA-Style pre-training with gradient-disentangled embedding sharing, 2023, arXiv preprint:2111.09543
Pengcheng He, Jianfeng Gao, and Weizhu Chen · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023, arXiv preprint:2309.00614
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Liangming Xia, Yijie Bai, Haiqin Weng, and Wenyuan Xu · 2024
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2024
Closest in time.
Lazy safety alignment for large language models against harmful fine-tuning, 2024, arXiv preprint:2405.18641
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2024
Closest in time.
Vaccine: Perturbation-aware alignment for large language model
Tiansheng Huang, Sihao Hu, and Ling Liu · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024, arXiv preprint:2401.05566
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Tuning language models by proxy, 2024, arXiv preprint:2401.08565
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith · 2024
Closest in time.
Keeping llms aligned after fine-tuning: The crucial role of prompt templates
Kaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu, Anirudh Goyal, and Sanjeev Arora · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024, arXiv preprint:2402.04249
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Fine-tuning can cripple foundation models; preserving features may be the solution, 2024
Jishnu Mukhoti, Yarin Gal, Philip Torr, and Puneet K. Dokania · 2024
Closest in time.
Smoothllm: Defending large language models against jailbreaking attacks, 2024, arXiv preprint:2310.03684
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas · 2024
Closest in time.
Immunization against harmful fine-tuning attacks, 2024, arXiv preprint:2402.16382
Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze, Jan Batzner, Hassan Sajjad, and Frank Rudzicz · 2024
Closest in time.
XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024, arXiv preprint:2308.01263
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy · 2024
Closest in time.
Fast yet effective machine unlearning
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, and Mohan Kankanhalli · 2024
Closest in time.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models, 2024, arXiv preprint:2306.11698
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson · 2024
Closest in time.
Open-source can be dangerous: On the vulnerability of value alignment in open-source LLMs, 2024
Jingwei Yi, Rui Ye, Qisi Chen, Bin Benjamin Zhu, Siheng Chen, Defu Lian, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2024
Closest in time.
A safety realignment framework via subspace-oriented model fusion for large language models
Xin Yi, Shunfan Zheng, Linlin Wang, Xiaoling Wang, and Liang He · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales · 2024
Closest in time.