Fetching the paper…
Reading the bibliography…
Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts.
Algorithmic fairness
Das, S., Stanton, R., and Wallace, N · 1941
Earlier work this paper cites.
Runaround
Asimov, I · 1942
Earlier work this paper cites.
The Human Use of Human Beings
Wiener, N · 1950
Earlier work this paper cites.
Intelligent machinery, a heretical theory
Turing, A · 1951
Earlier work this paper cites.
Speculations on perceptrons and other automata
Good, I. J · 1959
Earlier work this paper cites.
Gps, a program that simulates human thought
Newell, A. and Simon, H. A · 1961
Earlier work this paper cites.
Strips: A new approach to the application of theorem proving to problem solving
Fikes, R. E. and Nilsson, N. J · 1971
Earlier work this paper cites.
Problems of monetary management: the UK experience in papers in monetary economics
Goodhart, C · 1975
Earlier work this paper cites.
An introduction to fault tree analysis with emphasis on failure rate evaluation
Nieuwhof, G · 1975
Earlier work this paper cites.
An improved failures model for communicating processes
Brookes, S. D. and Roscoe, A. W · 1984
Earlier work this paper cites.
A theory of communicating sequential processes
Brookes, S. D., Hoare, C. A., and Roscoe, A. W · 1984
Earlier work this paper cites.
On the limitations of Markovian rewards to express multi-objective, risk-sensitive, and modal tasks
Skalse, J. and Abate, A · 1984
Earlier work this paper cites.
Afterword to Vernor Vinge’s novel,“True name
Minsky, M · 1986
Earlier work this paper cites.
Recognizing safety and liveness
Alpern, B. and Schneider, F. B · 1987
Earlier work this paper cites.
ind Children: The Future of Robot and Human Intelligence
Moravec, H · 1988
Earlier work this paper cites.
Expert system verification and validation: a survey and tutorial
O’Keefe, R. M. and O’Leary, D. E · 1993
Earlier work this paper cites.
The coming technological singularity: How to survive in the post-human era
Vinge, V · 1993
Earlier work this paper cites.
A logic for reasoning about time and reliability
Hansson, H. and Jonsson, B · 1994
Earlier work this paper cites.
An Introduction to Computational Learning Theory
Kearns, M. J. and Vazirani, U · 1994
Earlier work this paper cites.
Evaluating motion strategies under nondeterministic or probabilistic uncertainties in sensing and control
LaValle, S. M. and Hutchinson, S. A · 1996
Earlier work this paper cites.
Why the future doesn’t need us, 2000
Joy, B · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S · 2000
Earlier work this paper cites.
The Great Bridge: The Epic Story of the Building of the Brooklyn Bridge
McCullough, D · 2001
Earlier work this paper cites.
Creating friendly ai 1.0: The analysis and design of benevolent goal architectures
Yudkowsky, E · 2001
Earlier work this paper cites.
Existential risks: Analyzing human extinction scenarios and related hazards
Bostrom, N · 2002
Earlier work this paper cites.
Labeled rtdp: Improving the convergence of real-time dynamic programming
Bonet, B. and Geffner, H · 2003
Earlier work this paper cites.
Engineering Safety: Fundamentals, Techniques, And Applications
Dhillon, B · 2003
Earlier work this paper cites.
A gentle introduction to the universal algorithmic agent aixi, 2003
Hutter, M · 2003
Earlier work this paper cites.
Automated Planning: theory and practice
Ghallab, M., Nau, D., and Traverso, P · 2004
Earlier work this paper cites.
Bounded real-time dynamic programming: Rtdp with monotone upper bounds and performance guarantees
McMahan, H. B., Likhachev, M., and Gordon, G. J · 2005
Earlier work this paper cites.
The role of nondeterminism in model verification and validation
BenThacker, Andereson, C., Senseny, P., and Rodriguez, E · 2006
Earlier work this paper cites.
Sequential monte carlo samplers
Del Moral, P., Doucet, A., and Jasra, A · 2006
Earlier work this paper cites.
Blog: Probabilistic models with unknown objects
Milch, B., Marthi, B., Russell, S., Sontag, D., Ong, D. L., and Kolobov, A · 2007
Earlier work this paper cites.
Principles of Model Checking
Baier, C. and Katoen, J.-P · 2008
Earlier work this paper cites.
Solving pomdps: Rtdp-bel vs. point-based algorithms
Bonet, B. and Geffner, H · 2009
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Verification and control of hybrid systems: a symbolic approach
Tabuada, P · 2009
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Watanabe, S · 2009
Earlier work this paper cites.
Inverse optimal control with linearly-solvable MDPs
Dvijotham, K. and Todorov, E · 2010
Earlier work this paper cites.
A probabilistic model of theory formation
Kemp, C., Tenenbaum, J. B., Niyogi, S., and Griffiths, T. L · 2010
Earlier work this paper cites.
The gleamviz computational tool, a publicly available software to explore realistic epidemic spreading scenarios at the global scale
Broeck, W. V. d., Gioannini, C., Gonçalves, B., Quaggiotto, M., Colizza, V., and Vespignani, A · 2011
Earlier work this paper cites.
Thinking inside the box: Controlling and using an oracle ai
Armstrong, S., Sandberg, A., and Bostrom, N · 2012
Earlier work this paper cites.
Church: a language for generative models
Goodman, N., Mansinghka, V., Roy, D. M., Bonawitz, K., and Tenenbaum, J. B · 2012
Earlier work this paper cites.
Engineering a Safer World: Systems Thinking Applied to Safety
Leveson, N · 2012
Earlier work this paper cites.
Syntax-guided synthesis
Alur, R., Bodik, R., Juniwal, G., Martin, M. M. K., Raghothaman, M., Seshia, S. A., Singh, R., Solar-Lezama, A., Torlak, E., and Udupa, A · 2013
Earlier work this paper cites.
Halpern, J. Y. and Leung, S · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Earlier work this paper cites.
Bridge failure rates, consequences, and predictive trends
Cook, W · 2014
Earlier work this paper cites.
Hazard Analysis Techniques for System Safety
Ericson, C · 2015
Earlier work this paper cites.
Combining induction, deduction, and structure for verification and synthesis
Seshia, S. A · 2015
Earlier work this paper cites.
Corrigibility
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E · 2015
Earlier work this paper cites.
Coarse-to-fine sequential Monte Carlo for probabilistic programs
Stuhlmüller, A., Hawkins, R. X., Siddharth, N., and Goodman, N. D · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Actual Causality
Halpern, J. Y · 2016
Earlier work this paper cites.
Safety-constrained reinforcement learning for MDPs
Junges, S., Jansen, N., Dehnert, C., Topcu, U., and Katoen, J.-P · 2016
Earlier work this paper cites.
Inherent trade-offs in the fair determination of risk scores, 2016
Kleinberg, J., Mullainathan, S., and Raghavan, M · 2016
Earlier work this paper cites.
Automated theorem proving: A logical basis
Loveland, D. W · 2016
Earlier work this paper cites.
Safely interruptible agents
Orseau, L. and Armstrong, S · 2016
Earlier work this paper cites.
Towards Verified Artificial Intelligence
Seshia, S. A., Sadigh, D., and Sastry, S. S · 2016
Earlier work this paper cites.
Low impact artificial intelligences, 2017
Armstrong, S. and Levinstein, B · 2017
Earlier work this paper cites.
Program synthesis
Gulwani, S., Polozov, O., and Singh, R · 2017
Earlier work this paper cites.
The off-switch game, 2017
Hadfield-Menell, D., Dragan, A., Abbeel, P., and Russell, S · 2017
Earlier work this paper cites.
A Theory of Formal Synthesis via Inductive Learning
Jha, S. and Seshia, S. A · 2017
Earlier work this paper cites.
Unsupervised machine translation using monolingual corpora only
Lample, G., Conneau, A., Denoyer, L., and Ranzato, M · 2017
Earlier work this paper cites.
The hostile audience: The effect of access to broadband internet on partisan affect
Lelkes, Y., Sood, G., and Iyengar, S · 2017
Earlier work this paper cites.
Occam’s razor is insufficient to infer the preferences of irrational agents
Armstrong, S. and Mindermann, S · 2018
Earlier work this paper cites.
Good and safe uses of ai oracles, 2018
Armstrong, S. and O’Rorke, X · 2018
Earlier work this paper cites.
’indifference’ methods for managing agent rewards, 2018
Armstrong, S. and O’Rourke, X · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation, 2018
Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., Anderson, H., Roff, H., Allen, G. C., Steinhardt, J., Flynn, C., hEigeartaigh, S. O., Beard, S., Belfield, H., Farquhar, S., Lyle, C., Crootof, R., Evans, O., Page, M., Bryson, J., Yampolskiy, R., and Amodei, D · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Cited alongside, same era.
Counterexample-guided data augmentation
Dreossi, T., Ghosh, S., Yue, X., Keutzer, K., Sangiovanni-Vincentelli, A., and Seshia, S. A · 2018
Cited alongside, same era.
Learning linear temporal properties, 2018
Neider, D. and Gavran, I · 2018
Cited alongside, same era.
Formal specification for deep neural networks
Seshia, S. A., Desai, A., Dreossi, T., Fremont, D., Ghosh, S., Kim, E., Shivakumar, S., Vazquez-Chanlatte, M., and Yue, X · 2018
Cited alongside, same era.
Toward verified artificial intelligence
Seshia, S. A., Sadigh, D., and Sastry, S. S · 2022
Later among the works it cites.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals, 2022
Shah, R., Varma, V., Kumar, R., Phuong, M., Krakovna, V., Uesato, J., and Kenton, Z · 2022
Later among the works it cites.
Defining and Characterizing Reward Gaming
Skalse, J. M. V., Howe, N. H. R., Krasheninnikov, D., and Krueger, D · 2022
Later among the works it cites.
A complete criterion for value of information in soluble influence diagrams, 2022
van Merwijk, C., Carey, R., and Everitt, T · 2022
Later among the works it cites.
A brief review on algorithmic fairness
Wang, X., Zhang, Y., and Zhu, R · 2022
Later among the works it cites.
Autoformalization with large language models
Wu, Y., Jiang, A. Q., Li, W., Rabe, M., Staats, C., Jamnik, M., and Szegedy, C · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frenemies: How Social Media Polarizes America
Settle, J. E · 2018
Cited alongside, same era.
Life 3.0: Being human in the age of artificial intelligence
Tegmark, M · 2018
Cited alongside, same era.
Learning task specifications from demonstrations
Vazquez-Chanlatte, M., Jha, S., Tiwari, A., Ho, M. K., and Seshia, S. A · 2018
Cited alongside, same era.
Mathematical Theory of Bayesian Statistics
Watanabe, S · 2018
Cited alongside, same era.
Pyro: Deep universal probabilistic programming
Bingham, E., Chen, J. P., Jankowiak, M., Obermeyer, F., Pradhan, N., Karaletsos, T., Singh, R., Szerlip, P., Horsfall, P., and Goodman, N. D · 2019
Cited alongside, same era.
Disruptive innovations and disruptive assurance: Assuring machine learning and autonomy
Bloomfield, R., Khlaaf, H., Conmy, P. R., and Fletcher, G · 2019
Cited alongside, same era.
Liability, ethics, and culture-aware behavior specification using rulebooks
Censi, A., Slutsky, K., Wongpiromsarn, T., Yershov, D., Pendleton, S., Fu, J., and Frazzoli, E · 2019
Cited alongside, same era.
Later among the works it cites.
Adding native support for havoc in viper
Zhang, D · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability, 2022
Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N · 2022
Later among the works it cites.
Quantitative verification with neural networks
Abate, A., Edwards, A., Giacobbe, M., Punchihewa, H., and Roy, D · 2023
Later among the works it cites.
Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers, 2023
Alaga, J. and Schuett, J · 2023
Later among the works it cites.
Written testimony of Dario Amodei before the U.S. Senate Committee on the Judiciary, Subcommitee on Privacy, Technology, and the Law, 2023
Amodei, D · 2023
Later among the works it cites.
Anthropic’s responsible scaling policy, 2023
Anthropic · 2023
Later among the works it cites.
Written testimony of Yoshua Bengio before the U.S. Senate Committee on the Judiciary, Subcommitee on Privacy, Technology, and the Law, 2023
Bengio, Y · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., et al · 2023
Later among the works it cites.
Big ai can centralize decision-making and power, and that’s a problem
Brynjolfsson, E. and Ng, A · 2023
Later among the works it cites.
The hacking of chatgpt is just getting started, Apr 2023
Burgess, M · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback, 2023
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Later among the works it cites.
Joint bayesian inference of graphical structure and parameters with a single generative flow network, 2023
Deleu, T., Nishikawa-Toomey, M., Subramanian, J., Malkin, N., Charlin, L., and Bengio, Y · 2023
Later among the works it cites.
The open agency model, February 2023
Drexler, E · 2023
Later among the works it cites.
A general verification framework for dynamical and control models via certificate synthesis, 2023
Edwards, A., Peruffo, A., and Abate, A · 2023
Later among the works it cites.
Dreamcoder: growing generalizable, interpretable knowledge with wake–sleep bayesian program learning
Ellis, K., Wong, L., Nye, M., Sable-Meyer, M., Cary, L., Anaya Pozo, L., Hewitt, L., Solar-Lezama, A., and Tenenbaum, J. B · 2023
Later among the works it cites.
Baldur: Whole-proof generation and repair with large language models
First, E., Rabe, M., Ringer, T., and Brun, Y · 2023
Later among the works it cites.
AI governance scorecard and safety standards policy, 2023
FLI · 2023
Later among the works it cites.
Interpretability of machine learning: Recent advances and future prospects, 2023
Gao, L. and Guan, L · 2023
Later among the works it cites.
Bayes3d: fast learning and inference in structured generative models of 3d objects and scenes, 2023
Gothoskar, N., Ghavami, M., Li, E., Curtis, A., Noseworthy, M., Chung, K., Patton, B., Freeman, W. T., Tenenbaum, J. B., Klukas, M., and Mansinghka, V. K · 2023
Later among the works it cites.
Lilo: Learning interpretable libraries by compressing and documenting code
Grand, G., Wong, L., Bowers, M., Olausson, T. X., Liu, M., Tenenbaum, J. B., and Andreas, J · 2023
Later among the works it cites.
Certified reinforcement learning with logic guidance
Hasanbeig, H., Kroening, D., and Abate, A · 2023
Later among the works it cites.
An overview of catastrophic ai risks, 2023
Hendrycks, D., Mazeika, M., and Woodside, T · 2023
Later among the works it cites.
Goodhart’s Law and Machine Learning: A Structural Perspective
Hennessy, C. A. and Goodhart, C. A. E · 2023
Later among the works it cites.
‘the godfather of A.I.’ leaves Google and warns of danger ahead, May 2023
Hinton, G · 2023
Later among the works it cites.
TabPFN: A transformer that solves small tabular classification problems in a second, 2023
Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F · 2023
Later among the works it cites.
Gflownet-em for learning compositional latent variable models, 2023
Hu, E. J., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y · 2023
Later among the works it cites.
The consensus game: Language model generation via equilibrium search
Jacob, A. P., Shen, Y., Farina, G., and Andreas, J · 2023
Later among the works it cites.
Goodhart’s law in reinforcement learning, 2023
Karwowski, J., Hayman, O., Bai, X., Kiendlhofer, K., Griffin, C., and Skalse, J · 2023
Later among the works it cites.
Toward comprehensive risk assessments and assurance of ai-based systems
Khlaaf, H · 2023
Later among the works it cites.
Risk assessment at agi companies: A review of popular risk assessment techniques from other safety-critical industries, 2023
Koessler, L. and Schuett, J · 2023
Later among the works it cites.
Goal misgeneralization in deep reinforcement learning, 2023
Langosco, L., Koch, J., Sharkey, L., Pfau, J., Orseau, L., and Krueger, D · 2023
Later among the works it cites.
Ai safety on whose terms?, 2023
Lazar, S. and Nelson, A · 2023
Later among the works it cites.
Program Proofs
Leino, K. R. M · 2023
Later among the works it cites.
Beneficent intelligence: A capability approach to modeling benefit, assistance, and associated moral failures through ai systems, 2023
London, A. J. and Heidari, H · 2023
Later among the works it cites.
Proof Explorer - Home Page - Metamath, September 2023
Megill, N · 2023
Later among the works it cites.
Fairness, accountability, transparency, and ethics (fate) in artificial intelligence (ai) and higher education: A systematic review
Memarian, B. and Doleck, T · 2023
Later among the works it cites.
Reward Gaming in Conditional Text Generation, February 2023
Pang, R. Y., Padmakumar, V., Sellam, T., Parikh, A. P., and He, H · 2023
Later among the works it cites.
Written testimony of Stuart Russell before the U.S. Senate Committee on the Judiciary, Subcommitee on Privacy, Technology, and the Law, 2024
Russell, S · 2023
Later among the works it cites.
Sequential Monte Carlo learning for time series structure discovery
Saad, F., Patton, B., Hoffman, M. D., A. Saurous, R., and Mansinghka, V · 2023
Later among the works it cites.
Identifiability and generalizability in constrained inverse reinforcement learning, 2023
Schlaginhaufen, A. and Kamgarpour, M · 2023
Later among the works it cites.
Towards best practices in agi safety and governance: A survey of expert opinion, 2023
Schuett, J., Dreksler, N., Anderljung, M., McCaffary, D., Heim, L., Bluemke, E., and Garfinkel, B · 2023
Later among the works it cites.
Formal verification: an essential toolkit for modern VLSI design
Seligman, E., Schubert, T., and Kumar, M. A. K · 2023
Later among the works it cites.
Invariance in policy optimisation and partial identifiability in reward learning
Skalse, J. M. V., Farrugia-Roberts, M., Russell, S., Abate, A., and Gleave, A · 2023
Later among the works it cites.
Provably safe systems: the only path to controllable agi, 2023
Tegmark, M. and Omohundro, S · 2023
Later among the works it cites.
Compositional simulation-based analysis of AI-based autonomous systems for markovian specifications
Yalcinkaya, B., Torfah, H., Fremont, D. J., and Seshia, S. A · 2023
Later among the works it cites.
Towards a cautious scientist AI with convergent safety bounds, 2024
Bengio, Y · 2024
Closest in time.
Verified code transpilation with LLMs
Bhatia, S., Qiu, J., Hasabnis, N., Seshia, S. A., and Cheung, A · 2024
Closest in time.
Safeguarded ai: constructing guaranteed safety
’davidad’ Dalrymple, D · 2024
Closest in time.
Generating probabilistic scenario programs from natural language
Elmaaroufi, K., Shankar, D., Cismaru, A., Vazquez-Chanlatte, M., Sangiovanni-Vincentelli, A., Zaharia, M., and Seshia, S. A · 2024
Closest in time.
Cooperative inverse reinforcement learning, 2024
Hadfield-Menell, D., Dragan, A., Abbeel, P., and Russell, S · 2024
Closest in time.
Ai alignment: A comprehensive survey, 2024
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., Zeng, F., Ng, K. Y., Dai, J., Pan, X., O’Gara, A., Lei, Y., Xu, H., Tse, B., Fu, J., McAleer, S., Yang, Y., Wang, Y., Zhu, S.-C., Guo, Y., and Gao, W · 2024
Closest in time.
Compositional imprecise probability
Liell-Cock, J. and Staton, S · 2024
Closest in time.
Opening the ai black box: program synthesis via mechanistic interpretability
Michaud, E. J., Liao, I., Lad, V., Liu, Z., Mudide, A., Loughridge, C., Guo, Z. C., Kheirkhah, T. R., Vukelić, M., and Tegmark, M · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, J. H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V. H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M. B., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B · 2024
Closest in time.
Computing power and the governance of artificial intelligence, 2024
Sastry, G., Heim, L., Belfield, H., Anderljung, M., Brundage, M., Hazell, J., O’Keefe, C., Hadfield, G. K., Ngo, R., Pilz, K., Gor, G., Bluemke, E., Shoker, S., Egan, J., Trager, R. F., Avin, S., Weller, A., Bengio, Y., and Coyle, D · 2024
Closest in time.
Quantifying the sensitivity of inverse reinforcement learning to misspecification, 2024
Skalse, J. and Abate, A · 2024
Closest in time.
Starc: A general framework for quantifying differences between reward functions, 2024
Skalse, J., Farnik, L., Motwani, S. R., Jenner, E., Gleave, A., and Abate, A · 2024
Closest in time.
On the expressivity of objective-specification formalisms in reinforcement learning, 2024
Subramani, R., Williams, M., Heitmann, M., Holm, H., Griffin, C., and Skalse, J · 2024
Closest in time.
Tang, H., Key, D., and Ellis, K · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.
Probabilistic inference in language models via twisted sequential monte carlo
Zhao, S., Brekelmans, R., Makhzani, A., and Grosse, R · 2024
Closest in time.