Robust feature-level adversaries are interpretability tools
Original
Stephen Casper, Max Nadeau, and Gabriel Kreiman · 2021
Later among the works it cites.
Knowledge neurons in pretrained transformers
Original
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2021
Later among the works it cites.
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 2021
Later among the works it cites.
Naturalistic physical adversarial patch for object detectors
Yu-Chih-Tuan Hu, Bo-Han Kung, Daniel Stanley Tan, Jun-Cheng Chen, Kai-Lung Hua, and Wen-Huang Cheng · 2021
Later among the works it cites.
3db: A framework for debugging computer vision models
Original
Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, et al · 2021
Later among the works it cites.
Leveraging sparse linear layers for debuggable deep networks
Eric Wong, Shibani Santurkar, and Aleksander Madry · 2021
Later among the works it cites.
" real attackers don’t compute gradients": Bridging the gap between adversarial ml research and practice
Original
Giovanni Apruzzese, Hyrum S Anderson, Savino Dambra, David Freeman, Fabio Pierazzi, and Kevin A Roundy · 2022
Closest in time.
Domino: Discovering systematic errors with cross-modal embeddings
Original
Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Ré · 2022
Closest in time.
A survey on bias in visual datasets
Simone Fabbrizzi, Symeon Papadopoulos, Eirini Ntoutsi, and Ioannis Kompatsiaris · 2022
Closest in time.
Natural language descriptions of deep visual features
Original
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Closest in time.
Distilling model failures as directions in latent space, 2022
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Closest in time.
Locating and editing factual associations in gpt
Original
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Closest in time.
Red teaming language models with language models
Original
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Original
Tilman Räukur, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2022
Closest in time.
Chatgpt: Optimizing language models for dialogue, 2022
J Schulman, B Zoph, C Kim, J Hilton, J Menick, J Weng, JFC Uribe, L Fedus, L Metz, M Pokorny, et al · 2022
Closest in time.
Adversarial training for high-stakes reliability
Original
Daniel M Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, et al · 2022
Closest in time.
Benchmarking interpretability tools for deep neural networks
Original
Stephen Casper, Yuxiao Li, Jiawei Li, Tong Bu, Kevin Zhang, and Dylan Hadfield-Menell · 2023
Closest in time.
Benchmarking robustness to adversarial image obfuscations, 2023
Florian Stimberg, Ayan Chakrabarti, Chun-Ta Lu, Hussein Hazimeh, Otilia Stretcu, Wei Qiao, Yintao Liu, Merve Kaya, Cyrus Rashtchian, Ariel Fuxman, Mehmet Tek, and Sven Gowal · 2023
Closest in time.