Fetching the paper…
Reading the bibliography…
An agent trained within a closed system can master any desired capability, as long as the following three conditions hold: (a) it receives sufficiently informative and aligned feedback, (b) its coverage of experience/data is broad enough, and (c) it has sufficient capacity and resource.
Tractatus Logico-Philosophicus
Ludwig Wittgenstein · 1921
Earlier work this paper cites.
The machine age / In 1949, he imagined an age of robots
Norbert Wiener · 1949
Earlier work this paper cites.
Philosophical investigations
Ludwig Wittgenstein · 1953
Earlier work this paper cites.
Games people play: The psychology of human relationships , volume 2768
Eric Berne · 1968
Earlier work this paper cites.
The Rationalists
John Cottingham · 1988
Earlier work this paper cites.
A ‘self-referential’ weight matrix
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
Points… Interviews, 1974-1994
Jacques Derrida · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro et al · 1995
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive Levin search, and incremental self-improvement
Jürgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1997
Earlier work this paper cites.
Gödel machines: self-referential universal problem solvers making provably optimal self-improvements
Jürgen Schmidhuber · 2003
Earlier work this paper cites.
Metalearning
Tom Schaul and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Finite and infinite games
James Carse · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Deal or no deal? end-to-end learning for negotiation dialogues
Mike Lewis, Denis Yarats, Yann N Dauphin, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton · 2018
Earlier work this paper cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Cited alongside, same era.
Joel Z Leibo, Edward Hughes, Marc Lanctot, and Thore Graepel · 2019
Cited alongside, same era.
The bitter lesson
Richard S Sutton · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Games: Agency as art
C Thi Nguyen · 2020
Cited alongside, same era.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, Zheng Wen, et al · 2023
Later among the works it cites.
Levels of AGI: Operationalizing progress on the path to AGI
Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg · 2023
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Empirical design in reinforcement learning
Andrew Patterson, Samuel Neumann, Martha White, and Adam White · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miruna Pislar, David Szepesvari, Georg Ostrovski, Diana Borsa, and Tom Schaul · 2021
Cited alongside, same era.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S Sutton · 2021
Cited alongside, same era.
Language and culture internalization for human-like autotelic ai
Cédric Colas, Tristan Karch, Clément Moulin-Frier, and Pierre-Yves Oudeyer · 2022
Cited alongside, same era.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Diplomacy Team FAIR, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al · 2022
Cited alongside, same era.
Eliminating meta optimization through self-referential meta learning
Louis Kirsch and Jürgen Schmidhuber · 2022
Cited alongside, same era.
Exploration in deep reinforcement learning: A survey
Pawel Ladosz, Lilian Weng, Minwoo Kim, and Hyondong Oh · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Alexander Sasha Vezhnevets, John P Agapiou, Avia Aharon, Ron Ziv, Jayd Matyas, Edgar A Duéñez-Guzmán, William A Cunningham, Simon Osindero, Danny Karmon, and Joel Z Leibo · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
InterCode: Standardizing and benchmarking interactive coding with execution feedback
John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Three dogmas of reinforcement learning
David Abel, Mark K Ho, and Anna Harutyunyan · 2024
Closest in time.
AI achieves silver-medal standard solving International Mathematical Olympiad problems
AlphaProof and AlphaGeometry · 2024
Closest in time.
Does thought require sensory grounding? From pure thinkers to large language models
David J Chalmers · 2024
Closest in time.
Thousands of ai authors on the future of ai
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, and Jan Brauner · 2024
Closest in time.
Open-endedness is essential for artificial superhuman intelligence
Edward Hughes, Michael Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktäschel · 2024
Closest in time.
The AI scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Closest in time.
Learning to reason with LLMs
OpenAI et al · 2024
Closest in time.
Learning formal mathematics from intrinsic motivation
Gabriel Poesia, David Broman, Nick Haber, and Noah D Goodman · 2024
Closest in time.
Continual learning of large language models: A comprehensive survey
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, and Hao Wang · 2024
Closest in time.
GAVEL: Generating games via evolution and language models
Graham Todd, Alexander Padula, Matthew Stephenson, Éric Piette, Dennis JNJ Soemers, and Julian Togelius · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu · 2024
Closest in time.