Fetching the paper…
Reading the bibliography…
Large language models (LLMs) show inherent brittleness in their safety mechanisms, as evidenced by their susceptibility to jailbreaking and even non-malicious fine-tuning.
The Approximation of One Matrix by Another of Lower Rank
Eckart, C. and Young, G · 1936
Earlier work this paper cites.
Visualizing and Understanding Convolutional Networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
On Pixel-Wise Explanations for Non-linear Classifier Decisions by Layer-Wise Relevance Propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Striving for Simplicity: The All Convolutional Net
Springenberg, J., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2015
Earlier work this paper cites.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y · 2016
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
“Why Should I Trust You?” Explaining the Predictions of Any Classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Learning Structured Sparsity in Deep Neural Networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Earlier work this paper cites.
Channel Pruning for Accelerating Very Deep Neural Networks
He, Y., Zhang, X., and Sun, J · 2017
Earlier work this paper cites.
Pruning Filters for Efficient ConvNets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2017
Earlier work this paper cites.
Learning Efficient Convolutional Networks through Network Slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
Luo, J.-H., Wu, J., and Lin, W · 2017
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2017
Earlier work this paper cites.
Learning Important Features Through Propagating Activation Differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
What You Can Cram Into a Single Vector: Probing Sentence Embeddings for Linguistic Properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M · 2018
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models
Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, A., and Glass, J · 2019
Earlier work this paper cites.
Designing and Interpreting Probes With Control Tasks
Hewitt, J. and Liang, P · 2019
Earlier work this paper cites.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Lee, N., Ajanthan, T., and Torr, P · 2019
Earlier work this paper cites.
Are Sixteen Heads Really Better Than One?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Fine-Tuning Language Models From Human Preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Earlier work this paper cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Earlier work this paper cites.
A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
Tjoa, E. and Guan, C · 2020
Earlier work this paper cites.
Structured Pruning of Large Language Models
Wang, Z., Wohlwend, J., and Lei, T · 2020
Earlier work this paper cites.
Masking as an Efficient Alternative to Finetuning for Pretrained Language Models
Zhao, M., Lin, T., Mi, F., Jaggi, M., and Schütze, H · 2020
Earlier work this paper cites.
On the Pitfalls of Analyzing Individual Neurons in Language Models
Antverg, O. and Belinkov, Y · 2021
Cited alongside, same era.
A General Language Assistant as a Laboratory for Alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Cited alongside, same era.
A Survey on the Explainability of Supervised Machine Learning
Burkart, N. and Huber, M. F · 2021
Cited alongside, same era.
Low-Complexity Probing via Finding Subnetworks
Cao, S., Sanh, V., and Rush, A. M · 2021
Cited alongside, same era.
Parameter-efficient transfer learning with diff pruning
Guo, D., Rush, A. M., and Kim, Y · 2021
Cited alongside, same era.
Language Model Compression With Weighted Low-Rank Factorization
Hsu, Y.-C., Hua, T., Chang, S., Lou, Q., Shen, Y., and Jin, H · 2021
RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Later among the works it cites.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y · 2023
Later among the works it cites.
Mechanistic Mode Connectivity
Lubana, E. S., Bigelow, E. J., Dick, R. P., Krueger, D., and Tanaka, H · 2023
Later among the works it cites.
LLM-Pruner: On the Structural Pruning of Large Language Models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Can Neural Network Memorization Be Localized?
Maini, P., Mozer, M. C., Sedghi, H., Lipton, Z. C., Kolter, J. Z., and Zhang, C · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Block Pruning For Faster Transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Cited alongside, same era.
Circuit Component Reuse Across Tasks in Transformer Language Models
Merullo, J., Eickhoff, C., and Pavlick, E · 2021
Cited alongside, same era.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Probing Classifiers: Promises, Shortcomings, and Advances
Belinkov, Y · 2022
Cited alongside, same era.
Knowledge Neurons in Pretrained Transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Later among the works it cites.
GPT-4 Technical Report, 2023
OpenAI · 2023
Later among the works it cites.
Task-Specific Skill Localization in Fine-tuned Language Models
Panigrahi, A., Saunshi, N., Zhao, H., and Arora, S · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Sharma, P., Ash, J. T., and Misra, D · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., and Gan, C · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Function Vectors in Large Language Models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Jailbroken: How Does LLM Safety Training Fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Some things are more CRINGE than others: Preference Optimization with the Pairwise Cringe Loss
Xu, J., Lee, A., Sukhbaatar, S., and Weston, J · 2023
Later among the works it cites.
Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Low-Resource Languages Jailbreak GPT-4
Yong, Z. X., Menghini, C., and Bach, S · 2023
Later among the works it cites.
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
Yuan, Z., Shang, Y., Song, Y., Wu, Q., Yan, Y., and Sun, G · 2023
Later among the works it cites.
Removing RLHF Protections in GPT-4 via Fine-Tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Explainability for Large Language Models: A Survey
Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., and Du, M · 2023
Later among the works it cites.
AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Ashkboos, S., Croci, M. L., Nascimento, M. G. d., Hoefler, T., and Hensman, J · 2024
Closest in time.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2024
Closest in time.
Universal Neurons in GPT2 Language Models
Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D · 2024
Closest in time.
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R · 2024
Closest in time.
RAIN: Your Language Models Can Align Themselves without Finetuning
Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H · 2024
Closest in time.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2024
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2024
Closest in time.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2024
Closest in time.
RLCD: Reinforcement Learning from Contrastive Distillation for LM Alignment
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.