Fetching the paper…
Reading the bibliography…
Artificial intelligence (AI) has the potential to greatly improve society, but as with any powerful technology, it comes with heightened risks and responsibilities.
“The bases of social power.”, 1959
J.. French and Bertram. Raven · 1959
Earlier work this paper cites.
“Capitalism and Freedom”
Milton Friedman · 1963
Earlier work this paper cites.
“What you don’t have, can’t leak”
Trevor Kletz · 1978
Earlier work this paper cites.
“System Safety in Aircraft Acquisition”, 1984
F.. Frola and C.. Miller · 1984
Earlier work this paper cites.
“The Structure of Normative Ethics”
Shelly Kagan · 1992
Earlier work this paper cites.
“High Reliability Organizations: Unlikely, Demanding, and At Risk”
Todd La · 1996
Earlier work this paper cites.
“The Lack of A Priori Distinctions Between Learning Algorithms”
David. Wolpert · 1996
Earlier work this paper cites.
“How Complex Systems Fail”, 1998
Richard. Cook · 1998
Earlier work this paper cites.
“Normal Accidents: Living with High Risk Technologies”, Princeton paperbacks
C. Perrow · 1999
Earlier work this paper cites.
“Quadrennial Defense Review Report”, 2001
Department of Defense · 2001
Earlier work this paper cites.
“Existential risks: analyzing human extinction scenarios and related hazards”, 2002
Nick Bostrom · 2002
Earlier work this paper cites.
“The systems bible: the beginner’s guide to systems large and small”
John Gall · 2002
Earlier work this paper cites.
“Guide to emergency management and related terms, definitions, concepts, acronyms, organizations, programs, guidance, executive orders & legislation: A tutorial on emergency management, broadly defined, past and present”
B Blanchard · 2008
Earlier work this paper cites.
“Analysis of the 2007 Cyber Attacks Against Estonia from the Information Warfare Perspective”, 2008
Rain Ottis · 2008
Earlier work this paper cites.
“Moving Beyond Normal Accidents and High Reliability Organizations: A Systems Approach to Safety in Complex Systems”
Nancy. Leveson, Nicolas Dulac, Karen Marais and John. Carroll · 2009
Earlier work this paper cites.
“Global catastrophic risks”
Nick Bostrom and Milan Cirkovic · 2011
Earlier work this paper cites.
“For better or worse, benchmarks shape a field: technical perspective”
David. Patterson · 2012
Earlier work this paper cites.
“General Purpose Intelligence: Arguing the Orthogonality Thesis”, 2013
Stuart Armstrong · 2013
Earlier work this paper cites.
“Engineering a safer world: Systems thinking applied to safety”
Nancy Leveson · 2016
Earlier work this paper cites.
“The Off-Switch Game”
Dylan Hadfield-Menell, A. Dragan, P. Abbeel and Stuart. Russell · 2017
Earlier work this paper cites.
“Rate-Distortion Metrics for GAN”, 2017
David McAllester · 2017
Earlier work this paper cites.
“The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”
Miles Brundage, Shahar Avin, Jack Clark, H. Toner, P. Eckersley, Ben Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, Bobby Filar, H. Anderson, Heather Roff, Gregory. Allen, J. Steinhardt, Carrick Flynn, Se\’an\’O h\’Eigeartaigh, S. Beard, Haydn Belfield, Sebastian Farquhar, Clare Lyle, Rebecca Crootof, Owain Evans, Michael Page, Joanna Bryson, Roman Yampolskiy and Dario Amodei · 2018
Earlier work this paper cites.
“When will AI exceed human performance? Evidence from AI experts”
Katja Grace, John Salvatier, Allan Dafoe, Baobao Zhang and Owain Evans · 2018
Cited alongside, same era.
“Towards Deep Learning Models Resistant to Adversarial Attacks”
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras and Adrian Vladu · 2018
Cited alongside, same era.
“The Vulnerable World Hypothesis”
Nick Bostrom · 2019
Cited alongside, same era.
“The vulnerable world hypothesis”
Nick Bostrom · 2019
Cited alongside, same era.
“Using Pre-Training Can Improve Model Robustness and Uncertainty”
Dan Hendrycks, Kimin Lee and Mantas Mazeika · 2019
Cited alongside, same era.
“Deep Anomaly Detection with Outlier Exposure”
Dan Hendrycks, Mantas Mazeika and Thomas Dietterich · 2019
Cited alongside, same era.
“Statistical consequences of fat tails: Real world preasymptotics, epistemology, and applications”
Nassim Taleb · 2020
Later among the works it cites.
“Automating Cyber Attacks”, 2021
Ben Buchanan, John Bansemer, Dakota Cary, Jack Lucas and Micah Musser · 2021
Later among the works it cites.
“Evaluating large language models trained on code”
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph and Greg Brockman · 2021
Later among the works it cites.
“The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization”
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt and Justin Gilmer · 2021
Later among the works it cites.
“Aligning AI With Shared Human Values”
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song and Jacob Steinhardt · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty”
Dan Hendrycks, Mantas Mazeika, Saurav Kadavath and Dawn Song · 2019
Cited alongside, same era.
“Risks from Learned Optimization in Advanced Machine Learning Systems”
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse and Scott Garrabrant · 2019
Cited alongside, same era.
“The Bitter Lesson”, 2019
Richard Sutton · 2019
Cited alongside, same era.
“Neural cleanse: Identifying and mitigating backdoor attacks in neural networks”
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng and Ben Zhao · 2019
Cited alongside, same era.
“Building Secure and Reliable Systems: Best Practices for Designing, Implementing, and Maintaining Systems”
Heather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea and Adam Stubblefield · 2020
Cited alongside, same era.
“Destructive Cyber Operations and Machine Learning”, 2020
Dakota Cary and Daniel Cebul · 2020
Cited alongside, same era.
Later among the works it cites.
“Unsolved problems in ml safety”
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2021
Later among the works it cites.
“What Would Jiminy Cricket Do? Towards Agents That Behave Morally”
Dan Hendrycks, Mantas Mazeika, Andy Zou, Sahil Patel, Christine Zhu, Jesus Navarro, Dawn Song, Bo Li and Jacob Steinhardt · 2021
Later among the works it cites.
“The Parliamentary Approach to Moral Uncertainty”, 2021
Toby Newberry and Toby Ord · 2021
Later among the works it cites.
“Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets”
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin and Vedant Misra · 2021
Later among the works it cites.
“Nines of safety: a proposed unit of measurement of risk”, 2021
Terence Tao · 2021
Later among the works it cites.
“Highly accurate protein structure prediction for the human proteome”
Kathryn Tunyasuvunakool, Jonas Adler, Zachary Wu, Tim Green, Michal Zielinski, Augustin Z\’dek, Alex Bridgland, Andrew Cowie, Clemens Meyer and Agata Laydon · 2021
Later among the works it cites.
“Optimal Policies Tend To Seek Power”
Alexander Turner, Logan Smith, Rohin Shah, Andrew Critch and Prasad Tadepalli · 2021
Later among the works it cites.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”
Yushi Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort and Deep et al · 2022
Closest in time.
“Unsupervised Discovery of Latent Truth in Language Models”
Collin Burns, Haotian Ye, Dan Klein and Jacob Steinhardt · 2022
Closest in time.
“Is power-seeking AI an existential risk?”
Joseph Carlsmith · 2022
Closest in time.
“Jury Theorems”
Franz Dietrich and Kai Spiekermann · 2022
Closest in time.
“Execute Order 66: Targeted Data Poisoning for Reinforcement Learning”
Harrison Foley, Liam Fowl, Tom Goldstein and Gavin Taylor · 2022
Closest in time.
“Predictability and Surprise in Large Generative Models”
Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova DasSarma, T.. Henighan, Andy Jones, Nicholas Joseph, John Kernion, Benjamin Mann and Amanda et al · 2022
Closest in time.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike and Ryan. Lowe · 2022
Closest in time.
“Hierarchical Text-Conditional Image Generation with CLIP Latents”
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu and Mark Chen · 2022
Closest in time.
“Dual use of artificial-intelligence-powered drug discovery”
Fabio. Urbina, Filippa Lentzos, C\’edric Invernizzi and Sean Ekins · 2022
Closest in time.