Fetching the paper…
Reading the bibliography…
Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks.
Theory of Games and Economic Behavior (60th Anniversary Commemorative Edition)
John von Neumann, Oskar Morgenstern, and Ariel Rubinstein · 1944
Earlier work this paper cites.
The Foundations of Statistics
Leonard Savage · 1954
Earlier work this paper cites.
A formal theory of inductive inference. part i
R.J. Solomonoff · 1964
Earlier work this paper cites.
Learning from delayed rewards
Christopher Watkins · 1989
Earlier work this paper cites.
Allais Paradox , pages 3–9
Maurice Allais · 1990
Earlier work this paper cites.
Bolker-Jeffrey Expected Utility Theory and Axiomatic Utilitarianism
John Broome · 1990
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
A formal framework for agency and autonomy
Michael Luck and Mark d’Inverno · 1995
Earlier work this paper cites.
On tables of random numbers (reprinted from "sankhya: The indian journal of statistics", series a, vol. 25 part 4, 1963)
Andrei N. Kolmogorov · 1998
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity, 2000
Marcus Hutter · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
The Bayesian Choice: From Decision Theoretic Foundations to Computational Implementation
Christian Robert · 2007
Earlier work this paper cites.
Intentional systems theory
D. Dennett · 2009
Earlier work this paper cites.
Python 3 Reference Manual
Guido Van Rossum and Fred L. Drake · 2009
Earlier work this paper cites.
Dynamic Programming , volume 33
RICHARD BELLMAN and Stuart Dreyfus · 2010
Earlier work this paper cites.
A framework for goal generation and management
Marc Hanheide, Nick Hawes, Jeremy L. Wyatt, Moritz Göbelbecker, Michael Brenner, Kristoffer Sjöö, Alper Aydemir, Patric Jensfelt, Hendrik Zender, and Geert-Jan M. Kruijff · 2010
Earlier work this paper cites.
Hierarchical bayesian inverse reinforcement learning
Jaedeug Choi and Kee-Eung Kim · 2014
Earlier work this paper cites.
Formalizing preference utilitarianism in physical world models
Caspar Oesterheld · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris Maddison, et al · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks, 2016
Andreas Veit, Michael Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
Agents and devices: A relative definition of agency, 2018
Laurent Orseau, Simon McGregor McGill, and Shane Legg · 2018
Earlier work this paper cites.
Scikit-learn: Machine learning in python, 2018
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Andreas Müller, Joel Nothman, Gilles Louppe, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay · 2018
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems, 2019
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, et al · 2019
Cited alongside, same era.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Cited alongside, same era.
Parametrically retargetable decision-makers tend to seek power, 2022
Alexander Matt Turner and Prasad Tadepalli · 2022
Later among the works it cites.
Monte carlo tree search: a review of recent modifications and applications
Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk · 2022
Later among the works it cites.
Anthropic’s responsible scaling policy
Anthropic · 2023
Later among the works it cites.
Evaluating the historical value misspecification argument
Matthew Barnett · 2023
Later among the works it cites.
Features as the simplest factorization
Trenton Bricken, Joshua Batson, Adly Templeton, Adam Jermyn, Tom Henighan, and Chris Olah · 2023
Later among the works it cites.
Scheming ais: Will ais fake alignment during training in order to get power?, 2023
Joe Carlsmith · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Array programming with numpy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Cited alongside, same era.
Agi safety from first principles
Richard Ngo · 2020
Cited alongside, same era.
The Python Library Reference, release 3.8.2
Guido Van Rossum · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Perfectly secure steganography using minimum entropy coupling, 2023
Christian Schroeder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Foerster, and Martin Strohmeier · 2023
Later among the works it cites.
Discovering agents
Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt · 2023
Later among the works it cites.
Goal misgeneralization in deep reinforcement learning, 2023
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, Laurent Orseau, and David Krueger · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models, 2023
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Later among the works it cites.
Literature review on goal-directedness
Joe Collman Adam Shimi, Michele Campolo · 2024
Closest in time.
Counting arguments provide no evidence for ai doom
Nora Belrose and Quintin Pope · 2024
Closest in time.
tqdm: A fast, Extensible Progress Bar for Python and CLI, May 2024
Casper da Costa-Luis, Stephen Karl Larroque, Kyle Altendorf, Hadrien Mary, richardsheridan, Mikhail Korobov, Noam Yorav-Raphael, Ivan Ivanov, Marcel Bargull, Nishant Rodrigues, Guangshuo Chen, Mikhail Dektyarev, mjstevens777, Matthew D. Pagel, Martin Zugnoni, JC, CrazyPython, Charles Newey, Antony Lee, pgajdos, Todd, Staffan Malmgren, redbug312, Orivej Desh, Nikolay Nechaev, Michał Górny, Mike Boyle, Max Nordlund, MapleCCC, and Jack McCracken · 2024
Closest in time.
Understanding artificial agency
Le Kim Dung · 2024
Closest in time.
Model organisms of misalignment: The case for a new pillar of alignment research
Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez · 2024
Closest in time.
The learning-theoretic agenda: Status 2023
Vanessa Kosoy · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math, 2024
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Closest in time.
Shard theory: An overview
David Udell · 2024
Closest in time.
Coherence of caches and agents
John Wentworth · 2024
Closest in time.
Coherent decisions imply consistent utilities
Eliezer Yudkowsky · 2024
Closest in time.