Fetching the paper…
Reading the bibliography…
Is there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider estimating a context-dependent bound on the probability of violating a given safety specification.
Étude critique de la notion de collectif
Jean Ville · 1939
Earlier work this paper cites.
Application of the theory of martingales
J.L. Doob · 1949
Earlier work this paper cites.
On the asymptotic behavior of Bayes’ estimates in the discrete case
David A. Freedman · 1963
Earlier work this paper cites.
On the asymptotic behavior of Bayes estimates in the discrete case II
David A. Freedman · 1965
Earlier work this paper cites.
On Bayes procedures
Lorraine Schwartz · 1965
Earlier work this paper cites.
On the consistency of Bayes estimates (with discussion)
Persi Diaconis and David A. Freedman · 1986
Earlier work this paper cites.
The consistency of posterior distributions in nonparametric problems
Andrew Barron, Mark J. Schervish, and Larry Wasserman · 1999
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Quantitative risk management: concepts, techniques and tools-revised edition
Alexander J McNeil, Rüdiger Frey, and Paul Embrechts · 2015
Earlier work this paper cites.
A detailed treatment of Doob’s theorem
Jeffrey W. Miller · 2018
Earlier work this paper cites.
Pessimism about unknown unknowns inspires conservatism
Michael K Cohen and Marcus Hutter · 2020
Earlier work this paper cites.
Specification gaming: the flip side of AI ingenuity, 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
Simon Zhuang and Dylan Hadfield-Menell · 2020
Earlier work this paper cites.
Efficient (soft) Q-learning for text generation with limited good data
Han Guo, Bowen Tan, Zhengzhong Liu, Eric P. Xing, and Zhiting Hu · 2021
Cited alongside, same era.
Asymptotic normality, concentration, and coverage of generalized posteriors
Jeffrey W. Miller · 2021
Cited alongside, same era.
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Bayesian structure learning with generative flow networks
Tristan Deleu, António Góis, Chris Emezue, Mansi Rankawat, Simon Lacoste-Julien, Stefan Bauer, and Yoshua Bengio · 2022
Cited alongside, same era.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Towards a cautious scientist AI with convergent safety bounds, February 2024
Yoshua Bengio · 2024
Closest in time.
International Scientific Report on the Safety of Advanced AI
Yoshua Bengio, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, Shayne Longpre, et al · 2024
Closest in time.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al · 2024
Closest in time.
Amortizing intractable inference in large language models
Edward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, and Nikolay Malkin · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Defining and characterizing reward gaming
Joar Max Viktor Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Cited alongside, same era.
Joint Bayesian inference of graphical structure and parameters with a single generative flow network
Tristan Deleu, Mizu Nishikawa-Toomey, Jithendaraa Subramanian, Nikolay Malkin, Laurent Charlin, and Yoshua Bengio · 2023
Cited alongside, same era.
GFlowNet-EM for learning compositional latent variable models
Edward J. Hu, Nikolay Malkin, Moksh Jain, Katie Everett, Alexandros Graikos, and Yoshua Bengio · 2023
Cited alongside, same era.
Goodhart’s law in reinforcement learning
Jacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer, Charlie Griffin, and Joar Skalse · 2023
Cited alongside, same era.
Sequential Monte Carlo steering of large language models using probabilistic programs
Alexander K Lew, Tan Zhi-Xuan, Gabriel Grand, and Vikash K Mansinghka · 2023
Cited alongside, same era.
Reward Gaming in Conditional Text Generation
Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, and He He · 2023
Cited alongside, same era.
Training chain-of-thought via latent-variable inference
Du Phan, Matthew D. Hoffman, David Dohan, Sholto Douglas, Tuan Anh Le, Aaron Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, and Rif A. Saurous · 2023
Cited alongside, same era.
Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, and Nikolay Malkin · 2024
Closest in time.
STARC: A general framework for quantifying differences between reward functions
Joar Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, and Alessandro Abate · 2024
Closest in time.
Latent logic tree extraction for event sequence explanation from LLMs
Zitao Song, Chao Yang, Chaojie Wang, Bo An, and Shuang Li · 2024
Closest in time.
Amortizing intractable inference in diffusion models for vision, language, and control
Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, and Nikolay Malkin · 2024
Closest in time.
Flow of reasoning: Efficient training of LLM policy with divergent thinking
Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao, and Lianhui Qin · 2024
Closest in time.
Probabilistic inference in language models via twisted sequential Monte Carlo
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Baker Grosse · 2024
Closest in time.
PhyloGFN: Phylogenetic inference with generative flow networks
Ming Yang Zhou, Zichao Yan, Elliot Layne, Nikolay Malkin, Dinghuai Zhang, Moksh Jain, Mathieu Blanchette, and Yoshua Bengio · 2024
Closest in time.