Fetching the paper…
Reading the bibliography…
We consider the problem of low probability estimation: given a machine learning model and a formally-specified input distribution, how can we estimate the probability of a binary property of the model's output, even when that probability is too small to estimate by random sampling? This problem is motivated by the need to improve worst-case performance, which distribution shift can make much more likely.
Adversarial examples are not bugs, they are features, 2019
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 1905
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems, 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
The method of projections for finding the common point of convex sets
L.G. Gubin, Boris Polyak, and E.V. Raik · 1967
Earlier work this paper cites.
Analysis synthesis telephony based on the maximum likelihood method
F. Itakura and S. Saito · 1968
Earlier work this paper cites.
Adversarial training for large neural language models, 2020
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao · 2004
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
The metropolis-hastings algorithm, 2016
Christian P. Robert · 2016
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry · 2017
Earlier work this paper cites.
Models as approximations ii: A model-free theory of parametric regression, 2019
Andreas Buja, Lawrence Brown, Arun Kumar Kuchibhotla, Richard Berk, Ed George, and Linda Zhao · 2019
Earlier work this paper cites.
Transfer of adversarial robustness between perturbation types
Daniel Kang, Yi Sun, Tom Brown, Dan Hendrycks, and Jacob Steinhardt · 2019
Earlier work this paper cites.
A statistical approach to assessing neural network robustness, 2019
Stefan Webb, Tom Rainforth, Yee Whye Teh, and M. Pawan Kumar · 2019
Cited alongside, same era.
Recent advances in adversarial training for adversarial robustness, 2021
Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang · 2021
Cited alongside, same era.
Fudge: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Cited alongside, same era.
Formalizing the presumption of independence, 2022
Paul Christiano, Eric Neyman, and Mark Xu · 2022
Cited alongside, same era.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
A survey of controllable text generation using transformer-based pre-trained language models, 2023
Hanqing Zhang, Haolin Song, Shaoyu Li, Ming Zhou, and Dawei Song · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt · 2024
Closest in time.
Defending against unforeseen failure modes with latent adversarial training, 2024
Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell · 2024
Closest in time.
Analyzing probabilistic methods for evaluating agent capabilities, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Cited alongside, same era.
Natural Language Processing with Transformers: Building Language Applications with Hugging Face
Lewis Tunstall, Leandro von Werra, and Thomas Wolf · 2022
Cited alongside, same era.
Gaussian error linear units (gelus), 2023
Dan Hendrycks and Kevin Gimpel · 2023
Cited alongside, same era.
Sequential monte carlo steering of large language models using probabilistic programs, 2023
Alexander K. Lew, Tan Zhi-Xuan, Gabriel Grand, and Vikash K. Mansinghka · 2023
Cited alongside, same era.
Axel Højmark, Govind Pimpale, Arjun Panickssery, Marius Hobbhahn, and Jérémy Scheurer · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities, 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane · 2024
Closest in time.
Latent adversarial training improves robustness to persistent harmful behaviors in llms, 2024
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, and Stephen Casper · 2024
Closest in time.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
Estimating tail risk in neural networks
Mark Xu · 2024
Closest in time.
Probabilistic inference in language models via twisted sequential monte carlo, 2024
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Grosse · 2024
Closest in time.