Fetching the paper…
Reading the bibliography…
We are at the cusp of a transition from "learning from data" to "learning what data to learn from" as a central focus of artificial intelligence (AI) research.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Equilibrium points in n-person games
J. F. Nash et al · 1950
Earlier work this paper cites.
The theory of statistical decision
L. J. Savage · 1951
Earlier work this paper cites.
The origins of intelligence in children
J. Piaget · 1952
Earlier work this paper cites.
The construction of reality in the child
J. Piaget · 1954
Earlier work this paper cites.
Moore’s law
G. Moore · 1965
Earlier work this paper cites.
Dynamic programming
R. Bellman · 1966
Earlier work this paper cites.
Model predictive heuristic control: Applications to industrial processes
J. Richalet, A. Rault, J. Testud, and J. Papon · 1978
Earlier work this paper cites.
Mind in society: Development of higher psychological processes
L. S. Vygotsky and M. Cole · 1978
Earlier work this paper cites.
Search strategies of foraging animals
W. J. O’brien, H. I. Browman, and B. I. Evans · 1990
Earlier work this paper cites.
Apple tasting and nearly one-sided learning
D. P. Helmbold, N. Littlestone, and P. M. Long · 1992
Earlier work this paper cites.
Improving generalization with active learning
D. Cohn, L. Atlas, and R. Ladner · 1994
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
A. Robins · 1995
Earlier work this paper cites.
The philosophy of artificial life
M. A. Boden · 1996
Earlier work this paper cites.
Artificial life: An overview
C. G. Langton · 1997
Earlier work this paper cites.
A classification of long-term evolutionary dynamics
M. A. Bedau, E. Snyder, and N. H. Packard · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
R. M. French · 1999
Earlier work this paper cites.
Ants at work: how an insect society is organized
D. M. Gordon · 1999
Earlier work this paper cites.
Human knowledge and the infinite regress of reasons
P. D. Klein · 1999
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
S. Singh, T. Jaakkola, M. L. Littman, and C. Szepesvári · 2000
Earlier work this paper cites.
Interactive machine learning
J. A. Fails and D. R. Olsen Jr · 2003
Earlier work this paper cites.
Open-ended artificial evolution
R. K. Standish · 2003
Earlier work this paper cites.
Spatial representation of shelter locations in meerkats, suricata suricatta
M. B. Manser and M. B. Bell · 2004
Earlier work this paper cites.
Gaussian processes for machine learning
C. K. Williams and C. E. Rasmussen · 2006
Earlier work this paper cites.
Universal algorithmic intelligence: A mathematical top→ down approach
M. Hutter · 2007
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
P.-Y. Oudeyer, F. Kaplan, and V. V. Hafner · 2007
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
J. Schmidhuber · 2007
Earlier work this paper cites.
Video suggestion and discovery for youtube: taking random walks through the view graph
S. Baluja, R. Seth, D. Sivakumar, Y. Jing, J. Yagnik, S. Kumar, D. Ravichandran, and M. Aly · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
A. Der Kiureghian and O. Ditlevsen · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
S. Singh, R. L. Lewis, A. G. Barto, and J. Sorg · 2010
Earlier work this paper cites.
The cultural niche: Why social learning is essential for human adaptation
R. Boyd, P. J. Richerson, and J. Henrich · 2011
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel · 2011
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
J. Lehman and K. O. Stanley · 2011
Earlier work this paper cites.
A monte-carlo aixi approximation
J. Veness, K. S. Ng, M. Hutter, W. Uther, and D. Silver · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
J. Schmidhuber · 2013
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Identifying necessary conditions for open-ended evolution through the artificial life world of chromaria
L. Soros and K. Stanley · 2014
Earlier work this paper cites.
Handbook of research on the education of young children
B. Spodek and O. N. Saracho · 2014
Earlier work this paper cites.
The secret of our success
J. Henrich · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
J. Schulman, N. Heess, T. Weber, and P. Abbeel · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Why greatness cannot be planned: The myth of the objective
K. O. Stanley and J. Lehman · 2015
Earlier work this paper cites.
Learning to generate textual data
G. Bouchard, P. Stenetorp, and S. Riedel · 2016
Earlier work this paper cites.
Strategic classification
M. Hardt, N. Megiddo, C. Papadimitriou, and M. Wootters · 2016
Earlier work this paper cites.
Political polarization on twitter: Implications for the use of social media in digital governments
S. Hong and S. H. Kim · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Earlier work this paper cites.
Data programming: Creating large training sets, quickly
A. J. Ratner, C. M. De Sa, S. Wu, D. Selsam, and C. Ré · 2016
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Do we need more training data?
X. Zhu, C. Vondrick, C. C. Fowlkes, and D. Ramanan · 2016
Earlier work this paper cites.
Statistical biases in information retrieval metrics for recommender systems
A. Bellogín, P. Castells, and I. Cantador · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
A. Kendall and Y. Gal · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Cited alongside, same era.
Continual curiosity-driven skill acquisition from high-dimensional video inputs for humanoid robots
V. R. Kompella, M. Stollenga, M. Luciw, and J. Schmidhuber · 2017
Cited alongside, same era.
Grammar variational autoencoder
M. J. Kusner, B. Paige, and J. M. Hernández-Lobato · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. Oord, and R. Munos · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Sample-efficient optimization in the latent space of deep generative models via weighted retraining
A. Tripp, E. Daxberger, and J. M. Hernández-Lobato · 2020
Later among the works it cites.
Enhanced poet: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
R. Wang, J. Lehman, A. Rawal, J. Zhi, Y. Li, J. Clune, and K. Stanley · 2020
Later among the works it cites.
A survey of exploration methods in reinforcement learning
S. Amin, M. Gomrokchi, H. Satija, H. van Hoof, and D. Precup · 2021
Later among the works it cites.
Stewardship of global collective behavior
J. B. Bak-Coleman, M. Alfano, W. Barfuss, C. T. Bergstrom, M. A. Centeno, I. D. Couzin, J. F. Donges, M. Galesic, A. S. Gersick, J. Jacquet, et al · 2021
Later among the works it cites.
On statistical bias in active learning: How and when to fix it
S. Farquhar, Y. Gal, and T. Rainforth · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An introduction to decision theory
M. Peterson · 2017
Cited alongside, same era.
Snorkel: Fast training set generation for information extraction
A. J. Ratner, S. H. Bach, H. R. Ehrenberg, and C. Ré · 2017
Cited alongside, same era.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Cited alongside, same era.
Open-endedness: The last grand challenge you’ve never heard of
K. O. Stanley, J. Lehman, and L. Soros · 2017
Cited alongside, same era.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. Mankowitz, and S. Mannor · 2017
Cited alongside, same era.
Brax–a differentiable physics engine for large scale rigid body simulation
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
Evocraft: A new challenge for open-endedness
D. Grbic, R. B. Palm, E. Najarro, C. Glanois, and S. Risi · 2021
Later among the works it cites.
Environment generation for zero-shot compositional reinforcement learning
I. Gur, N. Jaques, Y. Miao, J. Choi, M. Tiwari, H. Lee, and A. Faust · 2021
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2021
Later among the works it cites.
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
E. Hüllermeier and W. Waegeman · 2021
Later among the works it cites.
Selecting data augmentation for simulating interventions
M. Ilse, J. M. Tomczak, and P. Forré · 2021
Later among the works it cites.
Generative models as a data source for multiview representation learning
A. Jahanian, X. Puig, Y. Tian, and P. Isola · 2021
Later among the works it cites.
Learning the truth from only one side of the story
H. Jiang, Q. Jiang, and A. Pacchiano · 2021
Later among the works it cites.
Replay-guided adversarial environment design
M. Jiang, M. Dennis, J. Parker-Holder, J. Foerster, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Prioritized level replay
M. Jiang, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
AI for full self-driving
A. Karpathy · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stenetorp, R. Jia, M. Bansal, C. Potts, and A. Williams · 2021
Later among the works it cites.
Variational diffusion models
D. Kingma, T. Salimans, B. Poole, and J. Ho · 2021
Later among the works it cites.
Generative interventions for causal learning
C. Mao, A. Cha, A. Gupta, H. Wang, J. Yang, and C. Vondrick · 2021
Later among the works it cites.
A graph placement methodology for fast chip design
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nazi, et al · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2021
Later among the works it cites.
Shaking the foundations: delusions in sequence models for interaction and control
P. A. Ortega, M. Kunesch, G. Delétang, T. Genewein, J. Grau-Moya, J. Veness, J. Buchli, J. Degrave, B. Piot, J. Perolat, et al · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
M. Paul, S. Ganguli, and G. K. Dziugaite · 2021
Later among the works it cites.
Megaverse: Simulating embodied agents at one million experiences per second
A. Petrenko, E. Wijmans, B. Shacklett, and V. Koltun · 2021
Later among the works it cites.
Prefixrl: Optimization of parallel prefix circuits using deep reinforcement learning
R. Roy, J. Raiman, N. Kant, I. Elkin, R. Kirby, M. Siu, S. Oberman, S. Godil, and B. Catanzaro · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
M. Samvelyan, R. Kirk, V. Kurin, J. Parker-Holder, M. Jiang, E. Hambro, F. Petroni, H. Küttler, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Maximum likelihood training of score-based diffusion models
Y. Song, C. Durkan, I. Murray, and S. Ermon · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, et al · 2021
Later among the works it cites.
Self-supervised learning with data augmentations provably isolates content from style
J. Von Kügelgen, Y. Sharma, L. Gresele, W. Brendel, B. Schölkopf, M. Besserve, and F. Locatello · 2021
Later among the works it cites.
Noveld: A simple yet effective exploration criterion
T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. E. Gonzalez, and Y. Tian · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Closest in time.
Play in predictive minds: A cognitive theory of play
M. M. Andersen, J. Kiverstein, M. Miller, and A. Roepstorff · 2022
Closest in time.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Closest in time.
Consumer-lending discrimination in the fintech era
R. Bartlett, A. Morse, R. Stanton, and N. Wallace · 2022
Closest in time.
Deep surrogate assisted generation of environments
V. Bhatt, B. Tjanaka, M. C. Fontaine, and S. Nikolaidis · 2022
Closest in time.
Data distributional properties drive emergent in-context learning in transformers
S. C. Chan, A. Santoro, A. K. Lampinen, J. X. Wang, A. Singh, P. H. Richemond, J. McClelland, S. DeepMind, and F. Hill · 2022
Closest in time.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas, et al · 2022
Closest in time.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Closest in time.
Collective intelligence for deep learning: A survey of recent developments
D. Ha and Y. Tang · 2022
Closest in time.
Exploration via elliptical episodic bonuses
M. Henaff, R. Raileanu, M. Jiang, and T. Rocktäschel · 2022
Closest in time.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Closest in time.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Closest in time.
Deepreta: An eta post-processing system at scale
X. Hu, T. Binaykiya, E. Frank, and O. Cirit · 2022
Closest in time.
Grounding aleatoric uncertainty in unsupervised environment design
M. Jiang, M. D. Dennis, J. Parker-Holder, A. Lupu, H. Kuttler, E. Grefenstette, T. Rocktäschel, and J. N. Foerster · 2022
Closest in time.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Y. LeCun · 2022
Closest in time.
Evolution through large models
J. Lehman, J. Gordon, S. Jain, K. Ndousse, C. Yeh, and K. O. Stanley · 2022
Closest in time.
Streamingqa: A benchmark for adaptation to new knowledge over time in question answering models
A. Liška, T. Kočiskỳ, E. Gribovskaya, T. Terzi, E. Sezener, D. Agrawal, C. d. M. d’Autume, T. Scholtes, M. Zaheer, S. Young, et al · 2022
Closest in time.
How to stay curious while avoiding noisy tvs using aleatoric uncertainty estimation
A. Mavor-Parker, K. Young, C. Barry, and L. Griffin · 2022
Closest in time.
Prioritized training on points that are learnable, worth learning, and not yet learnt
S. Mindermann, J. M. Brauner, M. T. Razzak, M. Sharma, A. Kirsch, W. Xu, B. Höltgen, A. N. Gomez, A. Morisot, S. Farquhar, et al · 2022
Closest in time.
Revisiting popularity and demographic biases in recommender evaluation and effectiveness
N. Neophytou, B. Mitra, and C. Stinson · 2022
Closest in time.
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2022
Closest in time.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Closest in time.
Evolving curricula with regret-based environment design
J. Parker-Holder, M. Jiang, M. Dennis, M. Samvelyan, J. Foerster, E. Grefenstette, and T. Rocktäschel · 2022
Closest in time.
Mastering the game of stratego with model-free multiagent reinforcement learning
J. Perolat, B. de Vylder, D. Hennes, E. Tarassov, F. Strub, V. de Boer, P. Muller, J. T. Connor, N. Burch, T. Anthony, et al · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Beyond neural scaling laws: beating power law scaling via data pruning
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. S. Morcos · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Closest in time.
LaMDA: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Closest in time.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Closest in time.
Deep surrogate assisted map-elites for automated hearthstone deckbuilding
Y. Zhang, M. C. Fontaine, A. K. Hoover, and S. Nikolaidis · 2022
Closest in time.