Fetching the paper…
Reading the bibliography…
Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data.
Algorithmic probability
M. Hutter, S. Legg, and P. M. B. Vitányi · 1941
Earlier work this paper cites.
Three models for the description of language
N. Chomsky · 1956
Earlier work this paper cites.
On a family of turing machines and the related programming language
C. Böhm · 1964
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Inductive reasoning and kolmogorov complexity
M. Li and P. M. Vitanyi · 1992
Earlier work this paper cites.
The context-tree weighting method: Basic properties
F. M. Willems, Y. M. Shtarkov, and T. J. Tjalkens · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Reflections on “the context tree weighting method: Basic properties”
F. Willems, Y. Shtarkov, and T. Tjalkens · 1997
Earlier work this paper cites.
The context-tree weighting method: Extensions
F. M. Willems · 1998
Earlier work this paper cites.
The speed prior: A new simplicity measure yielding near-optimal computable predictions
J. Schmidhuber · 2002
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
M. Hutter · 2004
Earlier work this paper cites.
On universal prediction and Bayesian confirmation
M. Hutter · 2007
Earlier work this paper cites.
Universal prediction of selected bits
T. Lattimore, M. Hutter, and V. Gavane · 2011
Earlier work this paper cites.
A philosophical treatise of universal induction
S. Rathmanner and M. Hutter · 2011
Earlier work this paper cites.
(Non-)equivalence of universal priors
I. Wood, P. Sunehag, and M. Hutter · 2011
Earlier work this paper cites.
Introduction to the Theory of Computation
M. Sipser · 2012
Earlier work this paper cites.
On ensemble techniques for aixi approximation
J. Veness, P. Sunehag, and M. Hutter · 2012
Earlier work this paper cites.
Principles of solomonoff induction and aixi
P. Sunehag and M. Hutter · 2013
Earlier work this paper cites.
(non-) equivalence of universal priors
I. Wood, P. Sunehag, and M. Hutter · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Intelligence as inference or forcing Occam on the world
P. Sunehag and M. Hutter · 2014
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
A. Joulin and T. Mikolov · 2015
Cited alongside, same era.
Loss bounds and time complexity for speed priors
D. Filan, J. Leike, and M. Hutter · 2016
Cited alongside, same era.
Recurrent neural networks as weighted language recognizers
Y. Chen, S. Gilroy, A. Maletti, J. May, and K. Knight · 2017
Pre-training without natural images
H. Kataoka, K. Okayasu, A. Matsumoto, E. Yamagata, R. Yamada, N. Inoue, A. Nakamura, and Y. Satoh · 2020
Later among the works it cites.
Meta-trained agents implement bayes-optimal agents
V. Mikulik, G. Delétang, T. McGrath, T. Genewein, M. Martic, S. Legg, and P. Ortega · 2020
Later among the works it cites.
A provably stable neural network turing machine
J. Stogin, A. Mali, and C. L. Giles · 2020
Later among the works it cites.
Meta-learning in neural networks: A survey
T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Hutter · 2017
Cited alongside, same era.
Smart augmentation learning an optimal data augmentation strategy
J. Lemley, S. Bazrafkan, and P. Corcoran · 2017
Cited alongside, same era.
The effectiveness of data augmentation in image classification using deep learning
L. Perez and J. Wang · 2017
Cited alongside, same era.
A generalized characterization of algorithmic probability
T. F. Sterkenburg · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Input–output maps are strongly biased towards simple outputs
K. Dingle, C. Q. Camargo, and A. A. Louis · 2018
Cited alongside, same era.
Neural networks and the chomsky hierarchy
G. Deletang, A. Ruoss, J. Grau-Moya, T. Genewein, L. K. Wenliang, E. Catt, C. Cundy, M. Hutter, S. Legg, J. Veness, et al · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2022
Later among the works it cites.
Language modeling is compression, 2023
G. Delétang, A. Ruoss, P.-A. Duquenne, E. Catt, T. Genewein, C. Mattern, J. Grau-Moya, L. K. Wenliang, M. Aitchison, L. Orseau, M. Hutter, and J. Veness · 2023
Later among the works it cites.
Memory-based meta-learning on non-stationary distributions
T. Genewein, G. Delétang, A. Ruoss, L. K. Wenliang, E. Catt, V. Dutordoir, J. Grau-Moya, L. Orseau, M. Hutter, and J. Veness · 2023
Later among the works it cites.
A theory of emergent in-context learning as implicit structure induction
M. Hahn and N. Goyal · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
B. Liu, J. T. Ash, S. Goel, A. Krishnamurthy, and C. Zhang · 2023
Later among the works it cites.
On the computational complexity and formal hierarchy of second order recurrent neural networks
A. Mali, A. Ororbia, D. Kifer, and L. Giles · 2023
Later among the works it cites.
Do deep neural networks have an inbuilt occam’s razor?
C. Mingard, H. Rees, G. Valle-Pérez, and A. A. Louis · 2023
Later among the works it cites.
Brainf*ck
U. Müller · 2023
Later among the works it cites.
X. Wang, W. Zhu, and W. Y. Wang · 2023
Later among the works it cites.
An Introduction to Universal Artificial Intelligence
E. Catt, D. Quarel, and M. Hutter · 2024
Closest in time.