Fetching the paper…
Reading the bibliography…
Artificial intelligence (AI) is interacting with people at an unprecedented scale, offering new avenues for immense positive impact, but also raising widespread concerns around the potential for individual and societal harm.
Differential games I
Rufus Isaacs · 1954
Earlier work this paper cites.
Some moral and technical consequences of automation
Norbert Wiener · 1960
Earlier work this paper cites.
The bellman equation for minimizing the maximum cost
EN Barron and H Ishii · 1989
Earlier work this paper cites.
Dynamic noncooperative game theory
Tamer Başar and Geert Jan Olsder · 1998
Earlier work this paper cites.
Conflict resolution for air traffic management: A study in multiagent hybrid systems
Claire Tomlin, George J Pappas, and Shankar Sastry · 1998
Earlier work this paper cites.
A game theoretic approach to controller design for hybrid systems
C.J. Tomlin, J. Lygeros, and S. Shankar Sastry · 2000
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
On reachability and minimum cost optimal control
John Lygeros · 2004
Earlier work this paper cites.
A time-dependent Hamilton-Jacobi formulation of reachable sets for continuous dynamic games
Ian Mitchell, Alex Bayen, and Claire J. Tomlin · 2005
Earlier work this paper cites.
Hierarchical, hybrid framework for collision avoidance algorithms in the national airspace
Michael Vitus and Claire Tomlin · 2008
Earlier work this paper cites.
Anomaly detection: A survey
Varun Chandola, Arindam Banerjee, and Vipin Kumar · 2009
Earlier work this paper cites.
Reachable set computation for uncertain time-varying linear systems
Matthias Althoff, Colas Le Guernic, and Bruce H Krogh · 2011
Earlier work this paper cites.
Hamilton-Jacobi Formulation for Reach–Avoid Differential Games
K. Margellos and J. Lygeros · 2011
Earlier work this paper cites.
On making robots understand safety: Embedding injury knowledge into control
Sami Haddadin, Simon Haddadin, Augusto Khoury, Tim Rokahr, Sven Parusel, Rainer Burgkart, Antonio Bicchi, and Alin Albu-Schäffer · 2012
Earlier work this paper cites.
Next generation airborne collision avoidance system
Mykel J Kochenderfer, Jessica E Holland, and James P Chryssanthacopoulos · 2012
Earlier work this paper cites.
A machine learning approach for real-time reachability analysis
Ross E Allen, Ashley A Clark, Joseph A Starek, and Marco Pavone · 2014
Earlier work this paper cites.
Online verification of automated road vehicles using reachability analysis
Matthias Althoff and John M Dolan · 2014
Earlier work this paper cites.
The scenario approach for stochastic model predictive control with bounds on closed-loop constraint violations
Georg Schildbach, Lorenzo Fagiano, Christoph Frei, and Manfred Morari · 2014
Earlier work this paper cites.
Reach-avoid problems with time-varying dynamics, targets and constraints
J. Fisac, M. Chen, C. J. Tomlin, and S. Sastry · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Robots and robotic devices – Collaborative robots
ISO 15066:2016 · 2016
Earlier work this paper cites.
Hamilton-Jacobi Reachability: A brief overview and recent advances
Somil Bansal, Mo Chen, Sylvia Herbert, and Claire J Tomlin · 2017
Earlier work this paper cites.
Provably safe motion of mobile robots in human environments
Stefan B Liu, Hendrik Roehm, Christian Heinzemann, Ingo Lütkebohle, Jens Oehlerking, and Matthias Althoff · 2017
Earlier work this paper cites.
Safe exploration algorithms for reinforcement learning controllers
Tommaso Mannucci, Erik-Jan van Kampen, Cornelis De Visser, and Qiping Chu · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Earlier work this paper cites.
Unsupervised anomaly detection with generative adversarial networks to guide marker discovery
Thomas Schlegl, Philipp Seeböck, Sebastian M Waldstein, Ursula Schmidt-Erfurth, and Georg Langs · 2017
Earlier work this paper cites.
Robust online motion planning via contraction theory and convex optimization
Sumeet Singh, Anirudha Majumdar, Jean-Jacques Slotine, and Marco Pavone · 2017
Earlier work this paper cites.
Performance assessment framework for robotic systems
Roger V. Bostelman, Joseph A. Falco, Marek Franaszek, and Kamel S. Saidi · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Robust, informative human-in-the-loop predictions via empirical reachable sets
Katherine Driggs-Campbell, Roy Dong, and Ruzena Bajcsy · 2018
Earlier work this paper cites.
World models
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Deep anomaly detection with outlier exposure
Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich · 2018
Earlier work this paper cites.
Generative modeling of multimodal multi-human behavior
Boris Ivanovic, Edward Schmerling, Karen Leung, and Marco Pavone · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Earlier work this paper cites.
Control barrier functions: Theory and applications
Aaron D Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada · 2019
Earlier work this paper cites.
Bridging Hamilton-Jacobi safety analysis and reinforcement learning
J. F. Fisac, N. F. Lugovoy, V. Rubies-Royo, S. Ghosh, and C. J. Tomlin · 2019
Earlier work this paper cites.
Formal specification and verification of autonomous robotic systems: A survey
Matt Luckcuck, Marie Farrell, Louise A Dennis, Clare Dixon, and Michael Fisher · 2019
Earlier work this paper cites.
Towards responsibility-sensitive safety of automated vehicles with reachable set analysis
Piotr F. Orzechowski, Kun Li, and Martin Lauer · 2019
Cited alongside, same era.
A classification-based approach for approximate reachability
Vicenç Rubies-Royo, David Fridovich-Keil, Sylvia Herbert, and Claire J Tomlin · 2019
Cited alongside, same era.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Cited alongside, same era.
Preferences implicit in the state of the world
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Zero-shot goal-directed dialogue via rl on imagined conversations
Joey Hong, Sergey Levine, and Anca Dragan · 2023
Later among the works it cites.
ISAACS: Iterative Soft Adversarial Actor-Critic for Safety
Kai-Chieh Hsu, Duy Phuong Nguyen, and Jaime Fernàndez Fisac · 2023
Later among the works it cites.
Deception game: Closing the safety–learning loop in interactive robot autonomy
Haimin Hu, Zixu Zhang, Kensuke Nakamura, Andrea Bajcsy, and Jaime F Fisac · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Later among the works it cites.
System two safety
Shane Legg · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrea Bajcsy, Somil Bansal, Ellis Ratner, Claire J Tomlin, and Anca D Dragan · 2020
Cited alongside, same era.
Guaranteed obstacle avoidance for multi-robot operations with limited actuation: A control barrier function approach
Yuxiao Chen, Andrew Singletary, and Aaron D Ames · 2020
Cited alongside, same era.
Overcoming the curse of dimensionality for some hamilton–jacobi partial differential equations via neural network architectures
Jérôme Darbon, Gabriel P Langlois, and Tingwei Meng · 2020
Cited alongside, same era.
Learning to control in power systems: Design and analysis guidelines for concrete safety problems
Roel Dobbe, Patricia Hidalgo-Gonzalez, Stavros Karagiannopoulos, Rodrigo Henriquez-Auba, Gabriela Hug, Duncan S Callaway, and Claire J Tomlin · 2020
Cited alongside, same era.
Drocc: Deep robust one-class classification
Sachin Goyal, Aditi Raghunathan, Moksh Jain, Harsha Vardhan Simhadri, and Prateek Jain · 2020
Cited alongside, same era.
Learning-based model predictive control: Toward safe learning in control
Lukas Hewing, Kim P Wabersich, Marcel Menner, and Melanie N Zeilinger · 2020
Cited alongside, same era.
Pac-bayes control: Learning policies that provably generalize to novel environments, 2020
Anirudha Majumdar, Alec Farid, and Anoopkumar Sonar · 2020
Cited alongside, same era.
Robust safe learning and control in an unknown environment: An uncertainty-separated control barrier function approach
Jiacheng Li, Qingchen Liu, Wanxin Jin, Jiahu Qin, and Sandra Hirche · 2023
Later among the works it cites.
Generating formal safety assurances for high-dimensional reachability
Albert Lin and Somil Bansal · 2023
Later among the works it cites.
Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Dan Hendrycks, and David Wagner · 2023
Later among the works it cites.
Nash learning from human feedback
Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, et al · 2023
Later among the works it cites.
Online update of safety assurances using confidence-based predictions
Kensuke Nakamura and Somil Bansal · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Motionlm: Multi-agent motion forecasting as language modeling
Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S Refaat, Rami Al-Rfou, and Benjamin Sapp · 2023
Later among the works it cites.
Data-driven safety filters: Hamilton-jacobi reachability, control barrier functions, and predictive methods for uncertain systems
Kim P Wabersich, Andrew J Taylor, Jason J Choi, Koushil Sreenath, Claire J Tomlin, Aaron D Ames, and Melanie N Zeilinger · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Fact sheet: President biden issues executive order on safe, secure, and trustworthy artificial intelligence, 2023
United States White House · 2023
Later among the works it cites.
Learning physically simulated tennis skills from broadcast videos
Haotian Zhang, Ye Yuan, Viktor Makoviychuk, Yunrong Guo, Sanja Fidler, Xue Bin Peng, and Kayvon Fatahalian · 2023
Later among the works it cites.
Guided conditional diffusion for controllable traffic simulation
Ziyuan Zhong, Davis Rempe, Danfei Xu, Yuxiao Chen, Sushant Veer, Tong Che, Baishakhi Ray, and Marco Pavone · 2023
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Closest in time.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Aligning llm agents by learning latent preference from user edits
Ge Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro, and Dipendra Misra · 2024
Closest in time.
Lin Guan, Yifan Zhou, Denis Liu, Yantian Zha, Heni Ben Amor, and Subbarao Kambhampati · 2024
Closest in time.
All the news from openai’s first developer conference
Alex Heath · 2024
Closest in time.
I pitted chatgpt against a real financial advisor to help me save for retirement—and the winner is clear
Coryanne Hicks · 2024
Closest in time.
The safety filter: A unified view of safety-critical control in autonomous systems
Kai-Chieh Hsu, Haimin Hu, and Jaime Fernández Fisac · 2024
Closest in time.
The world of air transport in 2019
International Civil Aviation Organization ICAO · 2024
Closest in time.
Boeing built deadly assumptions into 737 max, blind to a late design change
Jack Nicas, Natalie Kitroeff, David Gelles, and James Glanz · 2024
Closest in time.
New video of bay bridge 8-car crash shows tesla abruptly braking in ’self-driving’ mode
Dan Noyes · 2024
Closest in time.
Our approach to alignment research
OpenAI · 2024
Closest in time.
Dall-e now available without waitlist
OpenAI · 2024
Closest in time.
DALL·E 2 preview: Risks and limitations
OpenAI · 2024
Closest in time.
Building safe artificial intelligence: specification, robustness, and assurance
Pedro A. Ortega, Vishal Maini, and DeepMind Safety Team · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Closest in time.
A european approach to artificial intelligence
The European Union · 2024
Closest in time.
“he would still be here”: Man dies by suicide after talking with ai chatbot, widow says
Chloe Xiang · 2024
Closest in time.