Fetching the paper…
Reading the bibliography…
Large language models (LLMs) exhibit excellent ability to understand human languages, but do they also understand their own language that appears gibberish to us? In this work we delve into this question, aiming to uncover the mechanisms underlying such behavior in LLMs.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Rationalizing neural predictions
T. Lei, R. Barzilay, and T. Jaakkola · 2016
Earlier work this paper cites.
The limitations of deep learning in adversarial settings
N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami · 2016
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
N. Carlini and D. Wagner · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou · 2017
Earlier work this paper cites.
news-please: A generic news crawler and extractor
F. Hamborg, N. Meuschke, C. Breitinger, and B. Gipp · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
M. Alzantot, Y. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi · 2018
Earlier work this paper cites.
Physical adversarial examples for object detectors
D. Song, K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, F. Tramer, A. Prakash, and T. Kohno · 2018
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Earlier work this paper cites.
This email could save your life: Introducing the task of email subject line generation
R. Zhang and J. Tetreault · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
Lowkey: Leveraging adversarial attacks to protect social media users from facial recognition
V. Cherepanova, M. Goldblum, H. Foley, S. Duan, J. Dickerson, G. Taylor, and T. Goldstein · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
C. Guo, A. Sablayrolles, H. Jégou, and D. Kiela · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Cited alongside, same era.
Demystifying prompts in language models via perplexity estimation
H. Gonen, S. Iyer, T. Blevins, N. A. Smith, and L. Zettlemoyer · 2022
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein · 2023
Later among the works it cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, et al · 2023
Later among the works it cites.
Open sesame! universal black box jailbreaking of large language models
R. Lapid, R. Langberg, and M. Sipper · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Where do models go wrong? parameter-space saliency maps for explainability
R. Levin, M. Shu, E. Borgnia, F. Huang, M. Goldblum, and T. Goldstein · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Cited alongside, same era.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
F. Perez and I. Ribeiro · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Cited alongside, same era.
Gpt-4 technical report
O. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, and S. A. et al · 2023
Cited alongside, same era.
The internal state of an llm knows when its lying
A. Azaria and T. Mitchell · 2023
Cited alongside, same era.
T. Lieberum, M. Rahtz, J. Kramár, G. Irving, R. Shah, and V. Mikulik · 2023
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
X. Liu, N. Xu, M. Chen, and C. Xiao · 2023
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Later among the works it cites.
Large language models sensitivity to the order of options in multiple-choice questions
P. Pezeshkpour and E. Hruschka · 2023
Later among the works it cites.
Eliciting language model behaviors using reverse language models
J. Pfau, A. Infanger, A. Sheshadri, A. Panda, J. Michael, and C. Huebner · 2023
Later among the works it cites.
M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr · 2023
Later among the works it cites.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Later among the works it cites.
Llm lies: Hallucinations are not bugs, but features as adversarial examples
J.-Y. Yao, K.-P. Ning, Z.-H. Liu, M.-N. Ning, and L. Yuan · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Fast adversarial attacks on language models in one gpu minute
V. S. Sadasivan, S. Saha, G. Sriramanan, P. Kattakinda, A. Chegini, and S. Feizi · 2024
Closest in time.
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi · 2024
Closest in time.