Fetching the paper…
Reading the bibliography…
Poisoning of data sets is a potential security threat to large language models that can lead to backdoored models.
Mining and summarizing customer reviews. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining . 168–177
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011 · 2011
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. In The IEEE International Conference on Computer Vision (ICCV)
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. In Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29. Curran Associates, Inc
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Toxic Comment Classification Challenge
Cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017 · 2017
Earlier work this paper cites.
Learning to Generate Reviews and Discovering Sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Earlier work this paper cites.
Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks. In Research in Attacks, Intrusions, and Defenses , Michael Bailey, Thorsten Holz, Manolis Stamatogiannakis, and Sotiris Ioannidis (Eds.). Springer International Publishing, Cham, 273–294
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018 · 2018
Earlier work this paper cites.
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Learning Gender-Neutral Word Embeddings. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, Brussels, Belgium, 4847–4853
Jieyu Zhao, Yichao Zhou, Zeyu Li, Wei Wang, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019 · 2019
Earlier work this paper cites.
OpenWebText Corpus
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019 · 2019
Earlier work this paper cites.
Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019) , Rada Mihalcea, Ekaterina Shutova, Lun-Wei Ku, Kilian Evang, and Soujanya Poria (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 85–91
Anne Lauscher and Goran Glavaš. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization. In International Conference on Learning Representations (ICLR)
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Reverse engineering recurrent networks for sentiment classification reveals line attractor dynamics. In Advances in Neural Information Processing Systems , Vol. 32
Niru Maheswaranathan, Alex Williams, Matthew Golub, Surya Ganguli, and David Sussillo. 2019 · 2019
Earlier work this paper cites.
Black is to Criminal as Caucasian is to Police: Detecting and Removing Multiclass Bias in Word Embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 615–621
Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
Effective Dimensionality Reduction for Word Embeddings. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019) , Isabelle Augenstein, Spandana Gella, Sebastian Ruder, Katharina Kann, Burcu Can, Johannes Welbl, Alexis Conneau, Xiang Ren, and Marek Rei (Eds.). Association for Computational Linguistics, Florence, Italy, 235–243
Vikas Raunak, Vivek Gupta, and Florian Metze. 2019 · 2019
Earlier work this paper cites.
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. 2020 · 2020
Earlier work this paper cites.
Weight Poisoning Attacks on Pretrained Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) . 2793–2806
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Towards Debiasing Sentence Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 5502–5515
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020 · 2020
Earlier work this paper cites.
Interpreting GPT: The Logit Lens
Nostalgebraist. 2020 · 2020
Cited alongside, same era.
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 7237–7256
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Adversarial Machine Learning-Industry Perspectives. In 2020 IEEE Security and Privacy Workshops (SPW) . 69–75
Ram Shankar Siva Kumar, Magnus Nyström, John Lambert, Andrew Marshall, Mario Goertzel, Andi Comissoneru, Matt Swann, and Sharon Xia. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP): System Demonstrations . 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations (2022-09-29)
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2022 · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Red Teaming Language Models with Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (Abu Dhabi, United Arab Emirates, 2022-12). Association for Computational Linguistics, 3419–3448
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Backdoor Attacks on Self-Supervised Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2022). 13337–13346
Aniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, and Hamed Pirsiavash. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 · 2020
Cited alongside, same era.
Spinning Sequence-to-Sequence Models with Meta-Backdoors
Eugene Bagdasaryan and Vitaly Shmatikov. 2021 · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Unsolved problems in ML safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (2021-10-06)
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Cited alongside, same era.
Towards Understanding and Mitigating Social Biases in Language Models. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 6565–6576
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021 · 2021
Cited alongside, same era.
Concealed Data Poisoning Attacks on NLP Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Online, 2021-06). Association for Computational Linguistics, 139–150
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. In The Eleventh International Conference on Learning Representations (2022-09-29)
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Later among the works it cites.
Sniper Backdoor: Single Client Targeted Backdoor Attack in Federated Learning. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) (Raleigh, NC, USA, 2023-02). IEEE, 377–391
Gorka Abad, Servio Paguada, Oğuzhan Ersoy, Stjepan Picek, Víctor Julio Ramírez-Durán, and Aitor Urbieta. 2023 · 2023
Closest in time.
Venomave: Targeted Poisoning Against Speech Recognition. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) (Raleigh, NC, USA, 2023-02). IEEE, 404–417
Hojjat Aghakhani, Lea Schönherr, Thorsten Eisenhofer, Dorothea Kolossa, Thorsten Holz, Christopher Kruegel, and Giovanni Vigna. 2023 · 2023
Closest in time.
Identifying and Mitigating the Security Risks of Generative AI
Clark Barrett, Brad Boyd, Ellie Burzstein, Nicholas Carlini, Brad Chen, Jihye Choi, Amrita Roy Chowdhury, Mihai Christodorescu, Anupam Datta, Soheil Feizi, Kathleen Fisher, Tatsunori Hashimoto, Dan Hendrycks, Somesh Jha, Daniel Kang, Florian Kerschbaum, Eric Mitchell, John Mitchell, Zulfikar Ramzan, Khawaja Shams, Dawn Song, Ankur Taly, and Diyi Yang. 2023 · 2023
Closest in time.
Poisoning Web-Scale Training Datasets is Practical
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. 2023a · 2023
Closest in time.
SoK: Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks. In First IEEE Conference on Secure and Trustworthy Machine Learning
Stephen Casper, Tilman Rauker, Anson Ho, and Dylan Hadfield-Menell. 2023 · 2023
Closest in time.
Reducing Certified Regression to Certified Classification for General Poisoning Attacks. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) (Raleigh, NC, USA, 2023-02). IEEE, 484–523
Zayd Hammoudeh and Daniel Lowd. 2023 · 2023
Closest in time.
Understanding Transformer Memorization Recall Through Idioms. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (Dubrovnik, Croatia, 2023-05). Association for Computational Linguistics, 248–264
Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023 · 2023
Closest in time.
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Closest in time.
Backdoor Attacks on Time Series: A Generative Approach. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) (Raleigh, NC, USA, 2023-02). IEEE, 392–403
Yujing Jiang, Xingjun Ma, Sarah Monazam Erfani, and James Bailey. 2023 · 2023
Closest in time.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023 · 2023
Closest in time.
Activation Addition: Steering Language Models Without Optimization
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Closest in time.
Poisoning Language Models During Instruction Tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Closest in time.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
Declan Grabb, Max Lamparth, and Nina Vasan. 2024 · 2024
Closest in time.
Human vs. Machine: Language Models and Wargames
M. Lamparth et al · 2024
Closest in time.
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
J. P. Rivera et al · 2024
Closest in time.