Fetching the paper…
Reading the bibliography…
Adversarial examples pose a significant challenge to the robustness, reliability and alignment of deep neural networks.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Counterspeculation, auctions, and competitive sealed tenders
Robert B. Wilson · 1977
Earlier work this paper cites.
Problems of monetary management: The u.k. experience
Charles Goodhart · 1981
Earlier work this paper cites.
Cubic convolution interpolation for digital image processing
Robert G Keys · 1981
Earlier work this paper cites.
The laplacian pyramid as a compact image code
P. Burt and E. Adelson · 1983
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Foundations of vision
Brian A Wandell · 1995
Earlier work this paper cites.
Modelling the power spectra of natural images: Statistics and information
A van der Schaaf and J H van Hateren · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Balanced allocations
Yossi Azar, Andrei Z Broder, Anna R Karlin, and Eli Upfal · 1999
Earlier work this paper cites.
The power of two random choices: A survey of techniques and results
Michael Mitzenmacher, Andrea W. Richa, and Ramesh Sitaraman · 2001
Earlier work this paper cites.
The role of fixational eye movements in visual perception
Susana Martinez-Conde, Stephen L Macknik, and David H Hubel · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Training independent subnetworks for robust prediction, 2021
Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew M. Dai, and Dustin Tran · 2010
Earlier work this paper cites.
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy, and Surya Ganguli · 2010
Earlier work this paper cites.
Intriguing properties of neural networks, 2013
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Cited alongside, same era.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Ai safety gridworlds, 2017
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning, 2021
Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov · 2021
Later among the works it cites.
Should ensemble members be calibrated?, 2021
Xixin Wu and Mark Gales · 2021
Later among the works it cites.
Pixels still beat text: Attacking the openai clip model with text patches and adversarial pixel perturbations. 2021
Stanislav Fort · 2021
Later among the works it cites.
Adversarial examples for the openai clip in its zero-shot classification regime and their semantic generalization, jan 2021b
Stanislav Fort · 2021
Later among the works it cites.
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles, 2017
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Cited alongside, same era.
Adversarial attacks and defences: A survey, 2018
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization, 2018
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Adversarial examples that fool both computer vision and time-limited humans, 2018
Gamaleldin F. Elsayed, Shreya Shankar, Brian Cheung, Nicolas Papernot, Alex Kurakin, Ian Goodfellow, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift, 2019
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Cited alongside, same era.
Later among the works it cites.
Adversarial vulnerability of powerful near out-of-distribution detection, 2022
Stanislav Fort · 2022
Later among the works it cites.
Stanislav Fort, Ekin Dogus Cubuk, Surya Ganguli, and Samuel S. Schoenholz · 2022
Later among the works it cites.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon, 2022
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Scaling laws for adversarial attacks on language model activations, 2023
Stanislav Fort · 2023
Later among the works it cites.
Robust principles: Architectural design principles for adversarially robust cnns, 2023
ShengYun Peng, Weilin Xu, Cory Cornelius, Matthew Hull, Kevin Li, Rahul Duggal, Mansi Phute, Jason Martin, and Duen Horng Chau · 2023
Later among the works it cites.
Better diffusion models further improve adversarial training, 2023
Zekai Wang, Tianyu Pang, Chao Du, Min Lin, Weiwei Liu, and Shuicheng Yan · 2023
Later among the works it cites.
Decoupled kullback-leibler divergence loss, 2023
Jiequan Cui, Zhuotao Tian, Zhisheng Zhong, Xiaojuan Qi, Bei Yu, and Hanwang Zhang · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2023
Later among the works it cites.
Adversarial robustness limits via scaling-law and human-alignment studies, 2024
Brian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, and Bhavya Kailkhura · 2024
Closest in time.
When do universal image jailbreaks transfer between vision-language models?, 2024
Rylan Schaeffer, Dan Valentine, Luke Bailey, James Chua, Cristóbal Eyzaguirre, Zane Durante, Joe Benton, Brando Miranda, Henry Sleight, John Hughes, Rajashree Agrawal, Mrinank Sharma, Scott Emmons, Sanmi Koyejo, and Ethan Perez · 2024
Closest in time.