Fetching the paper…
Reading the bibliography…
Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose catastrophic risks.
“The Units of Selection”
Richard. Lewontin · 1970
Earlier work this paper cites.
“Stratospheric sink for chlorofluoromethanes: chlorine atomc-atalysed destruction of ozone”
Mario. Molina and F. Rowland · 1974
Earlier work this paper cites.
“Cooperation under the Security Dilemma”
Robert. Jervis · 1978
Earlier work this paper cites.
“Three Mile Island: a report to the commissioners and to the public. Volume I”, 1979
Mitchell Rogovin and George. Jr · 1979
Earlier work this paper cites.
“Assessing the impact of planned social change”
Donald Campbell · 1979
Earlier work this paper cites.
“Reckless Homicide?: Ford’s Pinto Trial”
Lee Strobel · 1980
Earlier work this paper cites.
“Grimshaw v. Ford Motor Co.”, 1981
1981
Earlier work this paper cites.
“Machines vs. Workers”
Charlotte Curtis · 1983
Earlier work this paper cites.
“Normal Accidents: Living with High-Risk Technologies”
Charles Perrow · 1984
Earlier work this paper cites.
“System Safety in Aircraft Acquisition”, 1984
F.R. Frola and C.O. Miller · 1984
Earlier work this paper cites.
“War, Presidents, and Public Opinion”, UPA book
J.E. Mueller · 1985
Earlier work this paper cites.
“The Making of the Atomic Bomb”
Richard Rhodes · 1986
Earlier work this paper cites.
“Selling Autos by Selling Safety”
Paul. Judge · 1990
Earlier work this paper cites.
“Asbestos: scientific developments and implications for public policy.”
Brooke. Mossman et al · 1990
Earlier work this paper cites.
“Working in Practice But Not in Theory: Theoretical Challenges of “High-Reliability Organizations””
T.. Laporte and Paula. Consolini · 1991
Earlier work this paper cites.
“The Chernobyl Accident: Updating of INSAG-1”, 1992
International Atomic Energy Agency · 1992
Earlier work this paper cites.
“Political Liberalism”
John Rawls · 1993
Earlier work this paper cites.
“Pale Blue Dot: A Vision of the Human Future in Space”
Carl Sagan · 1994
Earlier work this paper cites.
“Infanticide in Lions: Consequences and Counterstrategies”
Anne Pusey and Craig Packer · 1994
Earlier work this paper cites.
“The Sverdlovsk anthrax outbreak of 1979.”
Matthew Meselson et al · 1994
Earlier work this paper cites.
“The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA”
Diane Vaughan · 1996
Earlier work this paper cites.
“Risk management in a Dynamic Society: A Modeling Problem”
Jens Rasmussen · 1996
Earlier work this paper cites.
“Activation of the human brain by monetary reward”
G. Thut et al · 1997
Earlier work this paper cites.
“Aum Shinrikyo: once and future threat?”
Keith Olson · 1999
Earlier work this paper cites.
“Tobacco smoke carcinogens and lung cancer.”
Stephen. Hecht · 1999
Earlier work this paper cites.
“The Orbitofrontal Cortex and Reward”
Edmund. Rolls · 2000
Earlier work this paper cites.
“Ford 100: Defective Pinto Almost Took Ford’s Reputation With It”
Robert Sherefkin · 2003
Earlier work this paper cites.
“Lead neurotoxicity in children: basic mechanisms and clinical correlates.”
Theodore. Lidsky and Jay. Schneider · 2003
Earlier work this paper cites.
“Three Faces of Desire”, Philosophy of Mind Series
T. Schroeder · 2004
Earlier work this paper cites.
“The Bhopal disaster and its aftermath: a review”
Edward Broughton · 2005
Earlier work this paper cites.
“The ’Hittite plague’, an epidemic of tularemia and the first record of biological warfare.”
Siro Trevisanato · 2007
Earlier work this paper cites.
“Structural realism”
John Mearsheimer · 2007
Earlier work this paper cites.
“Inside the Twisted Mind of the Security Professional”
Bruce Schneier · 2008
Earlier work this paper cites.
“The Fourth Quadrant: A Map of the Limits of Statistics”
Nassim Taleb · 2008
Earlier work this paper cites.
“The changing economics of DNA synthesis” Number: 12 Publisher: Nature Publishing Group
Robert Carlson · 2009
Earlier work this paper cites.
“From Tail Fins to Hybrids: How Detroit Lost Its Dominance of the U.S. Auto Market”
Thomas. Klier · 2009
Earlier work this paper cites.
“Social Parasitism among Ants: A Review”
Alfred Buschinger · 2009
Earlier work this paper cites.
“The Diffusion of Military Power: Causes and Consequences for International Politics”
Michael Horowitz · 2010
Earlier work this paper cites.
“The dependence of viral RNA replication on co-opted host factors”
Peter. Nagy and Judit Pogany · 2011
Earlier work this paper cites.
“Thalidomide: the tragedy of birth defects and the effective treatment of disease.”
James. Kim and Anthony. Scialli · 2011
Earlier work this paper cites.
“Final Safety Culture Policy Statement”, Federal Register, 2011, pp. 34773
U.S. Commission · 2011
Earlier work this paper cites.
“Aum Shinrikyo: Insights into How Terrorists Develop Biological and Chemical Weapons”, 2012
Richard Danzig et al · 2012
Earlier work this paper cites.
“Facebook use predicts declines in subjective well-being in young adults”
Ethan Kross et al · 2013
Earlier work this paper cites.
“Intriguing properties of neural networks”
Christian Szegedy et al · 2013
Earlier work this paper cites.
“On the overwhelming importance of shaping the far future”, 2013
Nick Beckstead · 2013
Earlier work this paper cites.
“Meet MonsterMind, the NSA Bot That Could Wage Cyberwar Autonomously”
Kim Zetter · 2014
Earlier work this paper cites.
“The Matter of Heartbleed”
Zakir Durumeric et al · 2014
Earlier work this paper cites.
“Lessons Learned from the Fukushima Nuclear Accident for Improving Safety of U.S. Nuclear Plants”
National Council et al · 2014
Earlier work this paper cites.
“Air Force Swears: Our Nuke Launch Code Was Never ’00000000”’
Dan Lamothe · 2014
Cited alongside, same era.
“Human rights vs. robot rights: Forecasts from Japan”
Jennifer Robertson · 2014
Cited alongside, same era.
“The Possibility of an Ongoing Moral Catastrophe”
Evan. Williams · 2015
Cited alongside, same era.
“Introducing OpenAI”, 2015
Greg Brockman, Ilya Sutskever and OpenAI · 2015
Cited alongside, same era.
“Taxonomy of Pathways to Dangerous Artificial Intelligence”
Roman Yampolskiy · 2016
Cited alongside, same era.
“Engineering a safer world: Systems thinking applied to safety”
Nancy Leveson · 2016
Cited alongside, same era.
“Faulty reward functions in the wild”, 2016
“What are you optimizing for? Aligning Recommender Systems with Human Values”
Jonathan Stray et al · 2021
Later among the works it cites.
“Unsolved Problems in ML Safety”
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2021
Later among the works it cites.
“The Parliamentary Approach to Moral Uncertainty”, 2021
Toby Newberry and Toby Ord · 2021
Later among the works it cites.
“Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics”
Kevin. Esvelt · 2022
Later among the works it cites.
“Adherence to and Compliance with Arms Control, Nonproliferation, and Disarmament Agreements and Commitments”, 2022
U.S. of State · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dario Amodei and Jack Clark · 2016
Cited alongside, same era.
Dylan Hadfield-Menell, Anca. Dragan, P. Abbeel and Stuart. Russell · 2016
Cited alongside, same era.
In Oxford Reference , 2016
“Lyndon Baines Johnson” · 2016
Cited alongside, same era.
“Online Human-Bot Interactions: Detection, Estimation, and Characterization”
Onur Varol et al · 2017
Cited alongside, same era.
“The Flash Crash: High-Frequency Trading in an Electronic Market”
Andrei Kirilenko, Albert. Kyle, Mehrdad Samadi and Tugkan Tuzun · 2017
Cited alongside, same era.
“The Radium Girls: The Dark Story of America’s Shining Women”
Kate Moore · 2017
Cited alongside, same era.
“Dual use of artificial-intelligence-powered drug discovery”
Fabio. Urbina, Filippa Lentzos, Cédric Invernizzi and Sean Ekins · 2022
Later among the works it cites.
“It will be the greatest intellectual achievement of all time. An achievement of science, of engineering, and of the humanities, whose significance is beyond humanity, beyond life, beyond good and bad.”
Richard Sutton [@RichardSSutton] · 2022
Later among the works it cites.
“Structured access to AI capabilities: an emerging paradigm for safe AI deployment”
Toby Shevlane · 2022
Later among the works it cites.
“Artificial Intelligence Act: How the EU can take on the challenge posed by general-purpose AI systems”
Maximilian Gahntz and Claire Pershan · 2022
Later among the works it cites.
“Artificial intelligence and the offense–defense balance in cyber security”
Matteo. Bonfanti · 2022
Later among the works it cites.
“Adversarial Policies Beat Professional-Level Go AIs”
Tony Wang et al · 2022
Later among the works it cites.
“X-Risk Analysis for AI Research”
Dan Hendrycks and Mantas Mazeika · 2022
Later among the works it cites.
“The impact of chief risk officer appointments on firm risk and operational efficiency”
Huashan Li, Hugo.. Lam, William Ho and Andy.. Yeung · 2022
Later among the works it cites.
“The effects of reward misspecification: Mapping and mitigating misaligned models”
Alexander Pan, Kush Bhatia and Jacob Steinhardt · 2022
Later among the works it cites.
“Human-level play in the game of Diplomacy by combining language models with strategic reasoning”
Anton Bakhtin et al · 2022
Later among the works it cites.
“In-context Learning and Induction Heads”
Catherine Olsson et al · 2022
Later among the works it cites.
“Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance”, 2023
Sarah. Carter, Jaime. Yassif and Chris Isaac · 2023
Closest in time.
“Can large language models democratize access to dual-use biotechnology?”, 2023
Emily Soice et al · 2023
Closest in time.
“AI Succession”
Richard Sutton · 2023
Closest in time.
“Artificial Influence: An Analysis Of AI-Driven Persuasion”
Matthew Burtell and Thomas Woodside · 2023
Closest in time.
“What happens when your AI chatbot stops loving you back?”
Anna Tong · 2023
Closest in time.
“Sans ces conversations avec le chatbot Eliza, mon mari serait toujours là”
Pierre-François Lovens · 2023
Closest in time.
“Towards best practices in AGI safety and governance: A survey of expert opinion”, 2023
Jonas Schuett et al · 2023
Closest in time.
Yonadav Shavit · 2023
Closest in time.
“The Threat of Offensive AI to Organizations”
Yisroel Mirsky et al · 2023
Closest in time.
“Bing’s AI Is Threatening Users. That’s No Laughing Matter”
Billy Perrigo · 2023
Closest in time.
“In A.I. Race, Microsoft and Google Choose Speed Over Caution”
Nico Grant and Karen Weise · 2023
Closest in time.
“737 Max crashes: Boeing says not guilty to fraud charge”
Theo Leggett · 2023
Closest in time.
“Examples of AI Improving AI”, 2023
Thomas Woodside et al · 2023
Closest in time.
“Natural Selection Favors AIs over Humans”
Dan Hendrycks · 2023
Closest in time.
“The Darwinian Argument for Worrying About AI”
Dan Hendrycks · 2023
Closest in time.
“Anthropic’s $5B, 4-year plan to take on OpenAI”
Kyle Wiggers, Devin Coldewey and Manish Singh · 2023
Closest in time.
“Statement on AI Risk (“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”)”, 2023
Center for AI · 2023
Closest in time.
“Sparks of Artificial General Intelligence: Early experiments with GPT-4”
Sébastien Bubeck et al · 2023
Closest in time.
“Building a Culture of Safety for AI: Perspectives and Challenges”
David Manheim · 2023
Closest in time.
“Confronting Tech Power”, 2023
Amba Kak and Sarah West · 2023
Closest in time.
“AI Safety – Emerging Technology Observatory Research Almanac”, 2023
Center for Security and Emerging Technology · 2023
Closest in time.
“Dead rats, dopamine, performance metrics, and peacock tails: proxy failure is an inherent risk in goal-oriented systems”
Yohan. John, Leigh Caldwell, Dakota. McCoy and Oliver Braganza · 2023
Closest in time.
“Existential Risk from Power-Seeking AI”
Joseph Carlsmith · 2023
Closest in time.
“Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark.”
Alexander Pan et al · 2023
Closest in time.
“Benchmarking Neural Network Proxy Robustness to Optimization Pressure”, 2023
Andy Zou et al · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez and Sam Bowman · 2023
Closest in time.
“Discovering Latent Knowledge in Language Models Without Supervision”
Collin Burns, Haotian Ye, Dan Klein and Jacob Steinhardt · 2023
Closest in time.
“Representation engineering: Understanding and controlling the inner workings of neural networks”, 2023
Andy Zou et al · 2023
Closest in time.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang et al · 2023
Closest in time.
Jiashu Xu et al · 2023
Closest in time.
“LEACE: Perfect linear concept erasure in closed form”
Nora Belrose et al · 2023
Closest in time.