Fetching the paper…
Reading the bibliography…
Transformer based large language models with emergent capabilities are becoming increasingly ubiquitous in society.
Generating textual adversarial examples for deep learning models: A survey
Zhang, W. E., Sheng, Q. Z., and Alhazmi, A · 1901
Earlier work this paper cites.
What does BERT look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 1906
Earlier work this paper cites.
Visualizing and measuring the geometry of BERT
Coenen, A., Reif, E., Yuan, A., Kim, B., Pearce, A., Viégas, F. B., and Wattenberg, M · 1906
Earlier work this paper cites.
On the validity of self-attention as explanation in transformer models
Brunner, G., Liu, Y., Pascual, D., Richter, O., and Wattenhofer, R · 1908
Earlier work this paper cites.
Revealing the dark secrets of BERT
Kovaleva, O., Romanov, A., Rogers, A., and Rumshisky, A · 1908
Earlier work this paper cites.
Universal adversarial triggers for NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 1908
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Dynamic programming
Bellman, R · 1966
Earlier work this paper cites.
Attention flows: Analyzing and comparing attention mechanisms in language models
DeRose, J. F., Wang, J., and Berger, M · 2009
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2012
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y · 2016
Earlier work this paper cites.
Universal adversarial perturbations
Moosavi-Dezfooli, S., Fawzi, A., Fawzi, O., and Frossard, P · 2016
Cited alongside, same era.
What do neural machine translation models learn about morphology?
Belinkov, Y., Durrani, N., Dalvi, F., Sajjad, H., and Glass, J. R · 2017
Cited alongside, same era.
Hotflip: White-box adversarial examples for NLP
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2017
Cited alongside, same era.
Deep contextualized word representations, 2018
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
A discussion of ’adversarial examples are not bugs, they are features’
Engstrom, L., Gilmer, J., Goh, G., Hendrycks, D., Ilyas, A., Madry, A., Nakano, R., Nakkiran, P., Santurkar, S., Tran, B., Tsipras, D., and Wallace, E · 2019
The earth is flat and the sun is not a star: The susceptibility of gpt-2 to universal adversarial triggers
Heidenreich, H. S. and Williams, J. R · 2021
Later among the works it cites.
Universal adversarial attacks with natural triggers for text classification
Song, L., Yu, X., Peng, H.-T., and Narasimhan, K · 2021
Later among the works it cites.
Robustness of explanation methods for nlp models, 2022
Atmakuri, S., Chheda, T., Kandula, D., Yadav, N., Lee, T., and Tuinhof, H · 2022
Later among the works it cites.
Attention understands semantic relations
Chizhikova, A., Murzakhmetov, S., Serikov, O., Shavrina, T., and Burtsev, M · 2022
Later among the works it cites.
Locating and editing factual associations in gpt, 2022
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features, 2019
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Cited alongside, same era.
Open sesame: Getting inside bert’s linguistic knowledge, 2019
Lin, Y., Tan, Y. C., and Frank, R · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline, 2019
Tenney, I., Das, D., and Pavlick, E · 2019
Cited alongside, same era.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Wallat, J., Singh, J., and Anand, A · 2020
Cited alongside, same era.
Causal abstractions of neural networks, 2021
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Cited alongside, same era.
“that is a suspicious reaction!”: Interpreting logits variation to detect NLP adversarial attacks
Mosca, E., Agarwal, S., Ramí rez, J. R., and Groh, G · 2022
Later among the works it cites.
A mechanistic interpretability analysis of grokking
Nanda, N. and Lieberum, T · 2022
Later among the works it cites.
Minimal: Mining models for universal adversarial triggers
Singla, Y. K., Parekh, S., Singh, S., Chen, C., Krishnamurthy, B., and Shah, R. R · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Closest in time.