Fetching the paper…
Reading the bibliography…
Associative memory and probabilistic modeling are two fundamental topics in artificial intelligence.
Dispersion on a sphere
R. A. Fisher · 1953
Earlier work this paper cites.
Who belongs in the family?
R. L. Thorndike · 1953
Earlier work this paper cites.
Remarks on some nonparametric estimates of a density function
M. Rosenblatt · 1956
Earlier work this paper cites.
On estimation of a probability density function and mode
E. Parzen · 1962
Earlier work this paper cites.
Non-parametric estimation of a multivariate probability density
V. A. Epanechnikov · 1969
Earlier work this paper cites.
Ferguson distributions via pólya urn schemes
D. Blackwell and J. B. MacQueen · 1973
Earlier work this paper cites.
Mixtures of dirichlet processes with applications to bayesian nonparametric problems
C. E. Antoniak · 1974
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Analytic theory of the ground state properties of a spin glass. i. ising spin glass
F. Tanaka and S. Edwards · 1980
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
J. J. Hopfield · 1982
Earlier work this paper cites.
Neurons with graded response have collective computational properties like those of two-state neurons
J. J. Hopfield · 1984
Earlier work this paper cites.
Information capacity of the hopfield model
Y. Abu-Mostafa and J. S. Jacques · 1985
Earlier work this paper cites.
Exchangeability and related topics
D. J. Aldous, I. A. Ibragimov, J. Jacod, and D. J. Aldous · 1985
Earlier work this paper cites.
Saturation level of the hopfield model for neural network
A. Crisanti, D. J. Amit, and H. Gutfreund · 1986
Earlier work this paper cites.
Computing with neural circuits: A model
J. J. Hopfield and D. W. Tank · 1986
Earlier work this paper cites.
The capacity of the hopfield associative memory
R. McEliece, E. Posner, E. Rodemich, and S. Venkatesh · 1987
Earlier work this paper cites.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
P. J. Rousseeuw · 1987
Earlier work this paper cites.
The space of interactions in neural network models
E. Gardner · 1988
Earlier work this paper cites.
A reliable data-based bandwidth selection method for kernel density estimation
S. J. Sheather and M. C. Jones · 1991
Earlier work this paper cites.
Kernel smoothing
M. P. Wand and M. C. Jones · 1994
Earlier work this paper cites.
On convergence properties of the em algorithm for gaussian mixtures
L. Xu and M. I. Jordan · 1996
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
R. M. Neal and G. E. Hinton · 1998
Earlier work this paper cites.
Mdl principle for robust vector quantisation
H. Bischof, A. Leonardis, and A. Selb · 1999
Earlier work this paper cites.
X-means: Extending k-means with efficient estimation of the number of clusters
D. Pelleg, A. W. Moore, et al · 2000
Earlier work this paper cites.
Estimating the number of clusters in a data set via the gap statistic
R. Tibshirani, G. Walther, and T. Hastie · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Storage capacity of attractor neural networks with depressing synapses
J. J. Torres, L. Pantic, and H. J. Kappen · 2002
Earlier work this paper cites.
Learning the k in k-means
G. Hamerly and C. Elkan · 2003
Earlier work this paper cites.
Optimization with em and expectation-conjugate-gradient
R. Salakhutdinov, S. T. Roweis, and Z. Ghahramani · 2003
Earlier work this paper cites.
Finding the number of clusters in a dataset: An information-theoretic approach
C. A. Sugar and G. M. James · 2003
Cited alongside, same era.
Pattern recognition and machine learning , volume 4
C. M. Bishop and N. M. Nasrabadi · 2006
Cited alongside, same era.
Combinatorial stochastic processes: Ecole d’eté de probabilités de saint-flour xxxii-2002
J. Pitman · 2006
Cited alongside, same era.
The elements of statistical learning: data mining, inference, and prediction , volume 2
T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman · 2009
Cited alongside, same era.
Dirichlet process
Y. W. Teh et al · 2010
Cited alongside, same era.
On the equivalence of hopfield networks and boltzmann machines
A. Barra, A. Bernacchia, E. Santucci, and P. Contucci · 2012
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Q. Wang, B. Li, T. Xiao, J. Zhu, C. Li, D. F. Wong, and L. S. Chao · 2019
Later among the works it cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Later among the works it cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Memory engrams: Recalling the past and imagining the future
S. A. Josselyn and S. Tonegawa · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved contrastive divergence training of energy based models
Y. Du, S. Li, J. Tenenbaum, and I. Mordatch · 2012
Cited alongside, same era.
Revisiting k-means: new algorithms via bayesian nonparametrics
B. Kulis and M. I. Jordan · 2012
Cited alongside, same era.
Neurons are recruited to a memory trace based on relative neuronal excitability immediately before training
A. P. Yiu, V. Mercaldo, C. Yan, B. Richards, A. J. Rashid, H.-L. L. Hsiang, J. Pressey, V. Mahadevan, M. M. Tran, S. A. Kushner, et al · 2014
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Cited alongside, same era.
Dense associative memory for pattern recognition
D. Krotov and J. J. Hopfield · 2016
Cited alongside, same era.
D. Krotov and J. Hopfield · 2020
Later among the works it cites.
The role of neuronal excitability, allocation to an engram and memory linking in the behavioral generation of a false memory in mice
J. M. Lau, A. J. Rashid, A. D. Jacob, P. W. Frankland, D. L. Schacter, and S. A. Josselyn · 2020
Later among the works it cites.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
E. Nijkamp, M. Hill, T. Han, S.-C. Zhu, and Y. N. Wu · 2020
Later among the works it cites.
Overparameterized neural networks implement associative memory
A. Radhakrishnan, M. Belkin, and C. Uhler · 2020
Later among the works it cites.
Hopfield networks is all you need
H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlović, G. K. Sandve, et al · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2020
Later among the works it cites.
Attention approximates sparse distributed memory
T. Bricken and C. Pehlevan · 2021
Later among the works it cites.
Unsupervised learning of compositional energy concepts
Y. Du, S. Li, Y. Sharma, J. Tenenbaum, and I. Mordatch · 2021
Later among the works it cites.
Efficient online inference for nonparametric mixture models
R. Schaeffer, B. Bordelon, M. Khona, W. Pan, and I. R. Fiete · 2021
Later among the works it cites.
On the relationship between variational inference and auto-associative memory
L. Annabi, A. Pitti, and M. Quoy · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
S. Chan, A. Santoro, A. Lampinen, J. Wang, A. Singh, P. Richemond, J. McClelland, and F. Hill · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. S. Liang, and G. Valiant · 2022
Later among the works it cites.
Universal hopfield networks: A general framework for single-shot associative memory models
B. Millidge, T. Salvatori, Y. Song, T. Lukasiewicz, and R. Bogacz · 2022
Later among the works it cites.
Content addressable memory without catastrophic forgetting by heteroassociation with a fixed scaffold
S. Sharma, S. Chandra, and I. Fiete · 2022
Later among the works it cites.
In search of dispersed memories: Generative diffusion models are associative memory networks
L. Ambrogioni · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Later among the works it cites.
Examining the engram encoding specificity hypothesis in mice
J. H. Jung, Y. Wang, A. J. Mocle, T. Zhang, S. Köhler, P. W. Frankland, and S. A. Josselyn · 2023
Later among the works it cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Re, et al · 2023
Later among the works it cites.
The exponential capacity of dense associative memories
C. Lucibello and M. Mézard · 2023
Later among the works it cites.
End-to-end differentiable clustering with associative memories
B. Saha, D. Krotov, M. J. Zaki, and P. Ram · 2023
Later among the works it cites.
Associative memory under the probabilistic lens: Improved transformers & dynamic memory creation
R. Schaeffer, M. Khona, N. Zahedi, I. R. Fiete, A. Gromov, and S. Koyejo · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
S. Zhai, T. Likhomanenko, E. Littwin, D. Busbridge, J. Ramapuram, Y. Zhang, J. Gu, and J. M. Susskind · 2023
Later among the works it cites.
The emergence of clusters in self-attention dynamics
B. Geshkovski, C. Letrouit, Y. Polyanskiy, and P. Rigollet · 2024
Closest in time.