Fetching the paper…
Reading the bibliography…
The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Risks from learned optimization in advanced machine learning systems”
Evan Hubinger et al · 1906
Earlier work this paper cites.
“Truth and Probability” 1999 electronic edition
Frank. Ramsey · 1926
Earlier work this paper cites.
“Theory of Games and Economic Behavior (60th Anniversary Commemorative Edition)”
John von Neumann, Oskar Morgenstern and Ariel Rubinstein · 1944
Earlier work this paper cites.
“Foundations of Statistics”
L.. Savage · 1954
Earlier work this paper cites.
“A formal theory of inductive inference. Part I”
Ray Solomonoff · 1964
Earlier work this paper cites.
“Speculations concerning the first ultraintelligent machine”
Irving Good · 1966
Earlier work this paper cites.
“Problems of monetary management: the UK experience”
Charles Goodhart · 1984
Earlier work this paper cites.
“The case for motivated reasoning.”
Ziva Kunda · 1991
Earlier work this paper cites.
“Unskilled and unaware of it: How difficulties in recognizing one’s own incompetence lead to inflated self-assessments”
Justin Kruger and David Dunning · 1999
Earlier work this paper cites.
“Extinctions in near time”
Ross.E. MacPhee · 1999
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2001
Earlier work this paper cites.
“The evolutionary impact of invasive species”
H.. Mooney and E.. Cleland · 2001
Earlier work this paper cites.
“The AI-Box experiment” Accessed: 2025-02-06, Blog post: https://www.yudkowsky.net/singularity/aibox , 2002
Eliezer. Yudkowsky · 2002
Earlier work this paper cites.
“On the sample complexity of reinforcement learning”, 2003
Sham Kakade · 2003
Earlier work this paper cites.
“AI research considerations for human existential safety (ARCHES)”
Andrew Critch and David Krueger · 2006
Earlier work this paper cites.
“Geoengineering the climate: science, governance and uncertainty”
John Shepherd · 2009
Earlier work this paper cites.
“Thinking, Fast and Slow”
Daniel Kahneman · 2011
Earlier work this paper cites.
“A precautionary principle for dual use research in the life sciences”
Frida Kuhlau, Anna Höglund, Kathinka Evers and Stefan Eriksson · 2011
Earlier work this paper cites.
“A connection between score matching and denoising autoencoders”
Pascal Vincent · 2011
Earlier work this paper cites.
“Thinking inside the box: Controlling and using an oracle AI”
Stuart Armstrong, Anders Sandberg and Nick Bostrom · 2012
Earlier work this paper cites.
“Model-based machine learning”
Christopher. Bishop · 2012
Earlier work this paper cites.
“The superintelligent will: Motivation and instrumental rationality in advanced artificial agents”
Nick Bostrom · 2012
Earlier work this paper cites.
“Monte Carlo methods in Bayesian computation”
Ming-Hui Chen, Qi-Man Shao and Joseph Ibrahim · 2012
Earlier work this paper cites.
“Open problems in cooperative AI”
Allan Dafoe et al · 2012
Earlier work this paper cites.
“Representation learning: A review and new perspectives”
Yoshua Bengio, Aaron Courville and Pascal Vincent · 2013
Earlier work this paper cites.
“Better mixing via deep representations”
Yoshua Bengio, Grégoire Mesnil, Yann Dauphin and Salah Rifai · 2013
Earlier work this paper cites.
“Introduction to imprecise probabilities”
Thomas Augustin, Frank Coolen, Gert De and Matthias Troffaes · 2014
Earlier work this paper cites.
“Superintelligence: Paths, Dangers, Strategies”
Nick Bostrom · 2014
Earlier work this paper cites.
“Explaining and harnessing adversarial examples”
Ian Goodfellow, Jonathon Shlens and Christian Szegedy · 2014
Earlier work this paper cites.
“Weight uncertainty in neural network”
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu and Daan Wierstra · 2015
Earlier work this paper cites.
“The precautionary principle: definitions, applications and governance” Accessed: 2025-02-06, 2015
Didier Bourguignon · 2015
Earlier work this paper cites.
“Mutualism”
J.L. Bronstein · 2015
Earlier work this paper cites.
“Probabilistic machine learning and artificial intelligence”
Zoubin Ghahramani · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network”
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Deep Learning”
Yann LeCun, Yoshua Bengio and Geoffrey Hinton · 2015
Earlier work this paper cites.
“Regret minimization and related decision rules”, 2015
YinYee Leung · 2015
Earlier work this paper cites.
“Deep unsupervised learning using nonequilibrium thermodynamics”
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan and Surya Ganguli · 2015
Earlier work this paper cites.
“Transforming our world: the 2030 Agenda for Sustainable Development” Accessed: 2025-02-06, Webpage: https://sdgs.un.org/2030agenda , 2015
United Nations · 2015
Earlier work this paper cites.
“End to end learning for self-driving cars”
Mariusz Bojarski · 2016
Earlier work this paper cites.
“Faulty reward functions in the wild” Accessed: 2025-02-06, 2016
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
“Cooperative inverse reinforcement learning”
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel and Anca Dragan · 2016
Earlier work this paper cites.
“Human vs. computer Go: Review and prospect”
Chang-Shing Lee et al · 2016
Earlier work this paper cites.
“In two moves, AlphaGo and Lee Sedol redefined the future” Accessed: 2025-02-06
Cade Metz · 2016
Earlier work this paper cites.
“What does the universal prior actually look like?” Accessed: 2025-02-06, Blog post: https://ordinaryideas.wordpress.com/2016/11/30/what-does-the-universal-prior-actually-look-like/ , 2016
Paul Christiano · 2016
Earlier work this paper cites.
“Artificial intelligence: a modern approach”
Stuart Russell and Peter Norvig · 2016
Earlier work this paper cites.
“Mastering the game of Go with deep neural networks and tree search”
David Silver et al · 2016
Earlier work this paper cites.
“Good and safe uses of AI Oracles”
Stuart Armstrong and Xavier O’Rorke · 2017
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul Christiano et al · 2017
Earlier work this paper cites.
“Imitation learning: A survey of learning methods”
Ahmed Hussein, Mohamed Gaber, Eyad Elyan and Chrisina Jayne · 2017
Earlier work this paper cites.
“Public views of machine learning”, 2017
Ipsos · 2017
Earlier work this paper cites.
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan and Stuart Russell · 2017
Earlier work this paper cites.
“Elements of causal inference: foundations and learning algorithms”
Jonas Peters, Dominik Janzing and Bernhard Schölkopf · 2017
Earlier work this paper cites.
“Mastering the game of Go without human knowledge”
David Silver et al · 2017
Earlier work this paper cites.
“Mutual information neural estimation”
Mohamed Belghazi et al · 2018
Earlier work this paper cites.
“Bayesian Occam’s razor is a razor of the people”
Thomas Blanchard, Tania Lombrozo and Shaun Nichols · 2018
Earlier work this paper cites.
“Accelerated modern human–induced species losses: Entering the sixth mass extinction”
Gerardo Ceballos et al · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano and Dario Amodei · 2018
Earlier work this paper cites.
“The basic AI drives”
Stephen Omohundro · 2018
Earlier work this paper cites.
“Deep reinforcement learning for de novo drug design”
Mariya Popova, Olexandr Isayev and Alexander Tropsha · 2018
Earlier work this paper cites.
“A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play”
David Silver et al · 2018
Earlier work this paper cites.
“Reinforcement Learning: An Introduction”
Richard. Sutton and Andrew. Barto · 2018
Earlier work this paper cites.
“Life 3.0: Being human in the age of artificial intelligence”
Max Tegmark · 2018
Earlier work this paper cites.
“Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche”
Christophe Coupé, Yoon Oh, Dan Dediu and François Pellegrino · 2019
Earlier work this paper cites.
“Summary of the 2018 Department of Defense Artificial Intelligence Strategy” Accessed: 2025-02-06, 2019
Department of Defense · 2019
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Artificial intelligence in clinical and genomic diagnostics”
Raquel Dias and Ali Torkamani · 2019
Earlier work this paper cites.
“Incomplete contracting and AI alignment”
Dylan Hadfield-Menell and Gillian Hadfield · 2019
Earlier work this paper cites.
“On Variational Bounds of Mutual Information”
Ben Poole et al · 2019
Earlier work this paper cites.
“Human compatible: AI and the problem of control”
Stuart Russell · 2019
Earlier work this paper cites.
“A meta-transfer objective for learning to disentangle causal mechanisms”
Yoshua Bengio et al · 2020
Earlier work this paper cites.
“Bayesian neural networks: An introduction and survey”
Ethan Goan and Clinton Fookes · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models”
Jonathan Ho, Ajay Jain and Pieter Abbeel · 2020
Cited alongside, same era.
“Causal discovery from heterogeneous/nonstationary data”
Biwei Huang et al · 2020
Cited alongside, same era.
“Physics-informed machine learning: case studies for weather and climate modelling”
Karthik Kashinath et al · 2020
Cited alongside, same era.
“International evaluation of an AI system for breast cancer screening”
Scott McKinney et al · 2020
Cited alongside, same era.
“Performative prediction”
Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner and Moritz Hardt · 2020
Cited alongside, same era.
“Mastering Atari, Go, chess and shogi by planning with a learned model”
Julian Schrittwieser et al · 2020
Cited alongside, same era.
“Iterated denoising energy matching for sampling from Boltzmann densities”
Tara Akhound-Sadegh et al · 2024
Later among the works it cites.
“Position: stop making unscientific AGI performance claims”
Patrick Altmeyer, Andrew. Demetriou, Antony Bartlett and Cynthia.. Liem · 2024
Later among the works it cites.
“Machines of Loving Grace” Accessed: 2025-02-06, Blog post: https://darioamodei.com/machines-of-loving-grace , 2024
Dario Amodei · 2024
Later among the works it cites.
“Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku” Accessed: 2025-02-06, Webpage: https://www.anthropic.com/news/3-5-models-and-computer-use , 2024
Anthropic · 2024
Later among the works it cites.
“The Claude 3 Model Family: Opus, Sonnet, Haiku” Accessed: 2025-02-06, 2024
Anthropic · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news”
Cristian Vaccari and Andrew Chadwick · 2020
Cited alongside, same era.
“Self-driving cars: A survey”
Claudine Badue et al · 2021
Cited alongside, same era.
“Flow network based generative models for non-iterative diverse candidate generation”
Emmanuel Bengio et al · 2021
Cited alongside, same era.
“Eliciting latent knowledge: How to tell if your eyes deceive you” Accessed: 2025-02-06, 2021
Paul Christiano, Ajeya Cotra and Mark Xu · 2021
Cited alongside, same era.
“A Novel Estimator of Mutual Information for Learning to Disentangle Textual Representations”
Pierre Colombo, Pablo Piantanida and Chloé Clavel · 2021
Cited alongside, same era.
“The geometry of uncertainty”
Fabio Cuzzolin · 2021
Cited alongside, same era.
“The free world must prevail” Accessed: 2025-02-06, Blog post: https://situational-awareness.ai/the-free-world-must-prevail/ , 2024
Leopold Aschenbrenner · 2024
Later among the works it cites.
“Investigating Generalization Behaviours of Generative Flow Networks”
Lazar Atanackovic and Emmanuel Bengio · 2024
Later among the works it cites.
“Current state of LLM Risks and AI Guardrails”
Suriya Ayyamperumal and Limin Ge · 2024
Later among the works it cites.
“Can a Bayesian Oracle Prevent Harm from an Agent?”
Yoshua Bengio et al · 2024
Later among the works it cites.
“Mechanistic Interpretability for AI Safety – A Review”
Leonard Bereska and Efstratios Gavves · 2024
Later among the works it cites.
“Project Zero: from naptime to big sleep: using large language models to catch vulnerabilities in real-world code” Accessed: 2025-02-06, Blog post: https://googleprojectzero.blogspot.com/2024/10/from-naptime-to-big-sleep.html , 2024
Big Sleep Team · 2024
Later among the works it cites.
“The persuasive power of large language models”
Simon Breum et al · 2024
Later among the works it cites.
“The Oxford handbook of AI governance”
Justin Bullock et al · 2024
Later among the works it cites.
“NATO and Artificial Intelligence: Navigating the challenges and opportunities” Accessed: 2025-02-06, 2024
Sven Clement · 2024
Later among the works it cites.
“Regulating advanced artificial agents”
Michael Cohen et al · 2024
Later among the works it cites.
“Folk psychological attributions of consciousness to large language models”
Clara Colombatto and Stephen Fleming · 2024
Later among the works it cites.
“How much does it cost to train frontier AI models?” Accessed: 2025-02-06, 2024
Ben Cottier et al · 2024
Later among the works it cites.
“Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems”
David”davidad” Dalrymple et al · 2024
Later among the works it cites.
“Sycophancy to subterfuge: investigating reward-tampering in large language models”
Carson Denison et al · 2024
Later among the works it cites.
“LLM agents can autonomously exploit one-day vulnerabilities”
Richard Fang, Rohan Bindu, Akul Gupta and Daniel Kang · 2024
Later among the works it cites.
“LLM Agents can Autonomously Hack Websites”
Richard Fang et al · 2024
Later among the works it cites.
“System Design and Analysis” Accessed: 2025-02-06, 2024
Federal Aviation Administration · 2024
Later among the works it cites.
“The cognitive capabilities of generative AI: A comparative analysis with human benchmarks”
Isaac. Galatzer-Levy et al · 2024
Later among the works it cites.
“Jensen Huang wants Nvidia to have 100 million AI assistants” Accessed: 2025-02-06
Shubhangi Goel · 2024
Later among the works it cites.
“Build AI responsibly to benefit humanity” Accessed: 2025-02-06, Webpage: https://deepmind.google/about/ , 2024
Google DeepMind · 2024
Later among the works it cites.
“Thousands of AI authors on the future of AI”
Katja Grace et al · 2024
Later among the works it cites.
“Alignment faking in large language models”
Ryan Greenblatt et al · 2024
Later among the works it cites.
“AI Control: Improving Safety Despite Intentional Subversion”
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan and Fabien Roger · 2024
Later among the works it cites.
“Distorting the truth versus blatant lies: The effects of different degrees of deception in domestic and foreign political deepfakes”
Michael Hameleers, Toni.L.A. van Meer and Tom Dobber · 2024
Later among the works it cites.
“Rectifying reinforcement learning for reward matching”
Haoran He, Emmanuel Bengio, Qingpeng Cai and Ling Pan · 2024
Later among the works it cites.
“Rogue AIs” Accessed: 2025-02-06
Dan Hendrycks · 2024
Later among the works it cites.
“Amortizing intractable inference in large language models”
Edward Hu et al · 2024
Later among the works it cites.
“Data-efficient variational mutual information estimation via Bayesian self-consistency”
Desi. Ivanova, Marvin Schmitt and Stefan. Radev · 2024
Later among the works it cites.
“Uncovering deceptive tendencies in language models: A simulated company AI assistant”
Olli Järviniemi and Evan Hubinger · 2024
Later among the works it cites.
“SWE-Bench: Can language models resolve real-world Github issues?”
Carlos Jimenez et al · 2024
Later among the works it cites.
“Ant colony sampling with GFlowNets for combinatorial optimization”
Minsu Kim et al · 2024
Later among the works it cites.
“Neural general circulation models for weather and climate”
Dmitrii Kochkov et al · 2024
Later among the works it cites.
“Specification gaming: the flip side of AI ingenuity” Accessed: 2025-02-06, 2024
Victoria Krakovna et al · 2024
Later among the works it cites.
“Safe and trustworthy agents: an oxymoron?”
David Krueger · 2024
Later among the works it cites.
“A survey on model-based reinforcement learning”
Fan-Ming Luo et al · 2024
Later among the works it cites.
“An overview of the Gemini app” Accessed: 2025-02-06, 2024
James Manyika and Sissie Hsiao · 2024
Later among the works it cites.
“Large language model agents can coordinate beyond human scale”
Giordano Marzo, Claudio Castellano and David Garcia · 2024
Later among the works it cites.
“Artificial intelligence index report 2024” Accessed: 2025-02-06, 2024
Nestor Maslej et al · 2024
Later among the works it cites.
“Frontier models are capable of in-context scheming”
Alexander Meinke et al · 2024
Later among the works it cites.
“The alignment problem from a deep learning perspective”
Richard Ngo, Lawrence Chan and Sören Mindermann · 2024
Later among the works it cites.
“Learning to reason with LLMs” Accessed: 2025-02-06, Webpage: https://openai.com/index/learning-to-reason-with-llms/ , 2024
OpenAI · 2024
Later among the works it cites.
“MLE-bench: evaluating machine learning agents on machine learning engineering” Accessed: 2025-02-06, Webpage: https://openai.com/index/mle-bench/ , 2024
OpenAI · 2024
Later among the works it cites.
“OpenAI o1 System Card” Accessed: 2025-02-06, 2024
OpenAI · 2024
Later among the works it cites.
“AI deception: A survey of examples, risks, and potential solutions”
Peter Park et al · 2024
Later among the works it cites.
“Meta’s AI chief Yann LeCun on AGI, open-source, and AI risk” Accessed: 2025-02-06
Billy Perrigo · 2024
Later among the works it cites.
“A practical review of mechanistic interpretability for transformer-based language models”
Daking Rai et al · 2024
Later among the works it cites.
“Modern Bayesian experimental design”
Tom Rainforth, Adam Foster, Desi Ivanova and Freddie Bickford · 2024
Later among the works it cites.
“CtRL-Sim: reactive and controllable driving agents with offline reinforcement learning”
Luke Rowe et al · 2024
Later among the works it cites.
“Improved off-policy training of diffusion samplers”
Marcin Sendera et al · 2024
Later among the works it cites.
“The shutdown problem: an AI engineering puzzle for decision theorists”
Elliott Thornley · 2024
Later among the works it cites.
“Hacking CTFs with plain agents”
Rustem Turtayev, Artem Petrov, Dmitrii Volkov and Denis Volk · 2024
Later among the works it cites.
“Framework to Advance AI Governance and Risk Management in National Security” Accessed: 2025-02-06, White House publication: https://ai.gov/wp-content/uploads/2024/10/NSM-Framework-to-Advance-AI-Governance-and-Risk-Management-in-National-Security.pdf , 2024
US Government · 2024
Later among the works it cites.
“Amortizing Intractable Inference in Diffusion Models for Bayesian Inverse Problems” Accessed: 2025-02-06
Siddarth Venkatraman et al · 2024
Later among the works it cites.
“Position: will we run out of data? Limits of LLM scaling based on human-generated data”
Pablo Villalobos et al · 2024
Later among the works it cites.
“RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts”
Hjalmar Wijk et al · 2024
Later among the works it cites.
“International AI safety report 2025” Accessed: 2025-02-06, 2025
Yoshua Bengio et al · 2025
Closest in time.
“Robot Data Curation with Mutual Information Estimators”
Joey Hejna et al · 2025
Closest in time.
“I met the ‘godfathers of AI’ in Paris – here’s what they told me to really worry about” https://www.iaseai.org/conference/livestream , video Feb 7th 9:30 AM CET, at 57:51
Alexander Hurst · 2025
Closest in time.
“Causal discovery in astrophysics: Unraveling supermassive black hole and galaxy coevolution”
Zehao Jin et al · 2025
Closest in time.
“Adaptive teachers for amortized samplers”
Minsu Kim et al · 2025
Closest in time.
“Learning diverse attacks on large language models for robust red-teaming and safety tuning”
Seanie Lee et al · 2025
Closest in time.
“Announcing the Stargate Project” Accessed: 2025-02-06, Webpage: https://openai.com/index/announcing-the-stargate-project/ , 2025
OpenAI · 2025
Closest in time.
“Introducing ChatGPT search” Accessed: 2025-02-06, Webpage: https://openai.com/index/introducing-chatgpt-search/ , 2025
OpenAI · 2025
Closest in time.
“Introducing deep research” Accessed: 2025-02-06, Webpage: https://openai.com/index/introducing-deep-research/ , 2025
OpenAI · 2025
Closest in time.
“Introducing Operator” Accessed: 2025-02-06, Webpage: https://openai.com/index/introducing-operator/ , 2025
OpenAI · 2025
Closest in time.
“OpenAI o3-mini System Card” Accessed: 2025-02-06, 2025
OpenAI · 2025
Closest in time.
“Meta-Statistical Learning: Supervised Learning of Statistical Inference”
Maxime Peyrard and Kyunghyun Cho · 2025
Closest in time.