Fetching the paper…
Reading the bibliography…
Recent years have been marked with the fast-pace diversification and increasing ubiquity of machine learning applications.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Low-density parity-check codes
Robert Gallager · 1962
Earlier work this paper cites.
Some results on tchebycheffian spline functions
George Kimeldorf and Grace Wahba · 1971
Earlier work this paper cites.
Toward a mean field theory for spin glasses
Giorgio Parisi · 1979
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Order parameter for spin-glasses
Giorgio Parisi · 1983
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Sk model: The replica solution without replicas
M Mézard, G Parisi, and M Virasoro · 1987
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Virasoro · 1987
Earlier work this paper cites.
Optimal storage properties of neural network models
Elizabeth Gardner and Bernard Derrida · 1988
Earlier work this paper cites.
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
Judea Pearl · 1988
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
Elizabeth Gardner and Bernard Derrida · 1989
Earlier work this paper cites.
Storage capacity of memory networks with binary couplings
Werner Krauth and Marc Mézard · 1989
Earlier work this paper cites.
The space of interactions in neural networks: Gardner’s computation with the cavity method
Marc Mézard · 1989
Earlier work this paper cites.
Statistical mechanics of a multilayered neural network
E Barkai, D Hansel, and I Kanter · 1990
Earlier work this paper cites.
Neural networks and spin glasses
G Györgyi and N Tishby · 1990
Earlier work this paper cites.
First-order transition to perfect generalization in a neural network with binary synapses
Géza Györgyi · 1990
Earlier work this paper cites.
Learning from examples in large neural networks
Haim Sompolinsky, Naftali Tishby, and H Sebastian Seung · 1990
Earlier work this paper cites.
Calculation of the learning curve of bayes optimal classification algorithm for learning a perceptron with noise
Manfred Opper and David Haussler · 1991
Earlier work this paper cites.
Storage capacity and learning algorithms for two-layer neural networks
A Engel, HM Köhler, F Tschepke, H Vollmayr, and A Zippelius · 1992
Earlier work this paper cites.
Generalization properties of multilayered neural networks
German Mato and Nestor Parga · 1992
Earlier work this paper cites.
Generalization in a large committee machine
Henry Schwarze and John Hertz · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
Signature verification using a" siamese" time delay neural network
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger, and Roopak Shah · 1993
Earlier work this paper cites.
Learning a rule in a multilayer neural network
Henry Schwarze · 1993
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Domains of solutions and replica symmetry breaking in multilayer neural networks
Rémi Monasson and Dominic O’Kane · 1994
Earlier work this paper cites.
Learning and generalization in a two-layer neural network: The role of the vapnik-chervonvenkis dimension
Manfred Opper · 1994
Earlier work this paper cites.
Learning by on-line gradient descent
Michael Biehl and Holm Schwarze · 1995
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
On-line learning in the committee machine
Mauro Copelli and Nestor Caticha · 1995
Earlier work this paper cites.
Learning and generalization theories of large committee-machines
Rémi Monasson and Riccardo Zecchina · 1995
Earlier work this paper cites.
Weight space structure and internal representations: a direct approach to learning and generalization in multilayer neural networks
Rémi Monasson and Riccardo Zecchina · 1995
Earlier work this paper cites.
Statistical mechanics of learning: Generalization
Manfred Opper · 1995
Earlier work this paper cites.
On-line backpropagation in two-layered neural networks
Peter Riegler and Michael Biehl · 1995
Earlier work this paper cites.
Exact solution for on-line learning in multilayer neural networks
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A. Solla · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R.M. Neal · 1996
Earlier work this paper cites.
Mean field approach to bayes learning in feed-forward neural networks
Manfred Opper and Ole Winther · 1996
Earlier work this paper cites.
Learning with noise and regularizers in multilayer neural networks
David Saad and Sara Solla · 1996
Earlier work this paper cites.
Learning from minimum entropy queries in a large committee machine
Peter Sollich · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher K. I. Williams · 1996
Earlier work this paper cites.
Effect of batch learning in multilayer neural networks
Kenji Fukumizu · 1998
Earlier work this paper cites.
Belief propagation vs. tap for decoding corrupted messages
Yoshiyuki Kabashima and David Saad · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Mean field methods for classification with gaussian processes
Manfred Opper and Ole Winther · 1998
Earlier work this paper cites.
Statistical mechanics of support vector networks
Rainer Dietrich, Manfred Opper, and Haim Sompolinsky · 1999
Earlier work this paper cites.
On-line learning in neural networks
David Saad · 1999
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel · 2001
Earlier work this paper cites.
Universal learning curves of support vector machines
Manfred Opper and Robert Urbanczik · 2001
Earlier work this paper cites.
From naive mean field theory to the tap equations
Manfred Opper, Ole Winther, et al · 2001
Earlier work this paper cites.
A cdma multiuser detection algorithm on the basis of belief propagation
Yoshiyuki Kabashima · 2003
Earlier work this paper cites.
Superstatistics: theory and applications
Christian Beck · 2004
Earlier work this paper cites.
A bp-based algorithm for performing bayesian inference in large perceptron-type networks
Yoshiyuki Kabashima and Shinsuke Uda · 2004
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Kernel methods in machine learning
Thomas Hofmann, Bernhard Schölkopf, and Alexander J Smola · 2008
Earlier work this paper cites.
Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels
Yoshiyuki Kabashima · 2008
Earlier work this paper cites.
Modern coding theory
Tom Richardson and Ruediger Urbanke · 2008
Earlier work this paper cites.
Learning from correlated patterns by simple perceptrons
Takashi Shinzato and Yoshiyuki Kabashima · 2008
Earlier work this paper cites.
Perceptron capacity revisited: classification ability for correlated patterns
Takashi Shinzato and Yoshiyuki Kabashima · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Statistical mechanics of on-line learning
Michael Biehl, Nestor Caticha, and Peter Riegler · 2009
Earlier work this paper cites.
Message-passing algorithms for compressed sensing
David L Donoho, Arian Maleki, and Andrea Montanari · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Information, physics, and computation
Marc Mézard and Andrea Montanari · 2009
Earlier work this paper cites.
Message passing algorithms for compressed sensing: I. motivation and construction
David L Donoho, Arian Maleki, and Andrea Montanari · 2010
Earlier work this paper cites.
Statistical mechanics of compressed sensing
Surya Ganguli and Haim Sompolinsky · 2010
Earlier work this paper cites.
Statistical mechanical analysis of a typical reconstruction limit of compressed sensing
Yoshiyuki Kabashima, Tadashi Wadayama, and Toshiyuki Tanaka · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Mohsen Bayati and Andrea Montanari · 2011
Cited alongside, same era.
The lasso risk for gaussian matrices
Mohsen Bayati and Andrea Montanari · 2011
Cited alongside, same era.
Generalized approximate message passing for estimation with random linear mixing
Sundeep Rangan · 2011
Cited alongside, same era.
Probabilistic reconstruction in compressed sensing: algorithms, phase diagrams, and threshold achieving matrices
Florent Krzakala, Marc Mézard, Francois Sausset, Yifan Sun, and Lenka Zdeborová · 2012
Cited alongside, same era.
Statistical mechanics of complex neural systems and high dimensional data
Madhu Advani, Subhaneil Lahiri, and Surya Ganguli · 2013
Cited alongside, same era.
On robust regression with high-dimensional predictors
Noureddine El Karoui, Derek Bean, Peter J Bickel, Chinghway Lim, and Bin Yu · 2013
Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
Francesca Mignacco, Pierfrancesco Urbani, and Lenka Zdeborová · 2021
Later among the works it cites.
Analysis of feature learning in weight-tied autoencoders via the mean field lens
Phan-Minh Nguyen · 2021
Later among the works it cites.
Align, then memorise: the dynamics of learning with feedback alignment
Maria Refinetti, Stéphane d’Ascoli, Ruben Ohana, and Sebastian Goldt · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
Understanding the dynamics of gradient flow in overparameterized linear models
Salma Tarmoun, Guilherme Franca, Benjamin D Haeffele, and Rene Vidal · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
State evolution for general approximate message passing algorithms, with applications to spatial coupling
Adel Javanmard and Andrea Montanari · 2013
Cited alongside, same era.
The squared-error of generalized lasso: A precise analysis
Samet Oymak, Christos Thrampoulidis, and Babak Hassibi · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Origin of the computational hardness for learning with binary synapses
Haiping Huang and Yoshiyuki Kabashima · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
The influence of learning rule on representation dynamics in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
Exact learning dynamics of deep linear networks with prior knowledge
Lukas Braun, Clémentine Dominé, James Fitzgerald, and Andrew Saxe · 2022
Later among the works it cites.
Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Later among the works it cites.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Later among the works it cites.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2022
Later among the works it cites.
A precise high-dimensional asymptotic theory for boosting and minimum- ℓ \ell 1-norm interpolated classifiers
Tengyuan Liang and Pragya Sur · 2022
Later among the works it cites.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Later among the works it cites.
Fluctuations, bias, variance & ensemble of learners: Exact asymptotics for convex losses in high-dimension
Bruno Loureiro, Cédric Gerbelot, Maria Refinetti, Gabriele Sicuro, and Florent Krzakala · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
The effective noise of stochastic gradient descent
Francesca Mignacco and Pierfrancesco Urbani · 2022
Later among the works it cites.
Universality of empirical risk minimization
Andrea Montanari and Basil N Saeed · 2022
Later among the works it cites.
Reverend bayes on inference engines: A distributed hierarchical approach
Judea Pearl · 2022
Later among the works it cites.
The dynamics of representation learning in shallow, non-linear autoencoders
Maria Refinetti and Sebastian Goldt · 2022
Later among the works it cites.
Solvable model for inheriting the regularization through knowledge distillation
Luca Saglietti and Lenka Zdeborová · 2022
Later among the works it cites.
Precise learning curves and higher-order scalings for dot-product kernel regression
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington · 2022
Later among the works it cites.
Contrasting random and learned features in deep bayesian linear regression
Jacob Zavatone-Veth, Cengiz Pehlevan, and William L. Tong · 2022
Later among the works it cites.
High-dimensional robust regression under heavy-tailed data: Asymptotics and universality
Urte Adomaityte, Leonardo Defilippis, Bruno Loureiro, and Gabriele Sicuro · 2023
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Later among the works it cites.
Precise asymptotic analysis of deep random feature models
David Bosch, Ashkan Panahi, and Babak Hassibi · 2023
Later among the works it cites.
Fundamental limits of overparametrized shallow neural networks for supervised learning
Francesco Camilli, Daria Tieplova, and Jean Barbier · 2023
Later among the works it cites.
Theoretical characterization of uncertainty in high-dimensional linear classification
Lucas Clarté, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
Learning curves for the multi-class teacher–student perceptron
Elisabetta Cornacchia, Francesca Mignacco, Rodrigo Veiga, Cédric Gerbelot, Bruno Loureiro, and Lenka Zdeborová · 2023
Later among the works it cites.
Recent applications of dynamical mean-field methods
Leticia F Cugliandolo · 2023
Later among the works it cites.
Bayes-optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
Error scaling laws for kernel classification under source and capacity conditions
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
How two-layer neural networks learn, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
Neural networks: From the perceptron to deep nets
Marylou Gabrié, Surya Ganguli, Carlo Lucibello, and Riccardo Zecchina · 2023
Later among the works it cites.
Bayesian interpolation with deep linear networks
Boris Hanin, Alexander Zlokapa, and Alexander Zlokapa · 2023
Later among the works it cites.
Dataset size dependence of rate-distortion curve and threshold of posterior collapse in linear vae
Yuma Ichikawa and Koji Hukushima · 2023
Later among the works it cites.
High-dimensional manifold of solutions in neural networks: insights from statistical physics
Enrico M Malatesta · 2023
Later among the works it cites.
A theory of non-linear feature learning with one gradient step in two-layer neural networks
Behrad Moniri, Donghwan Lee, Hamed Hassani, and Edgar Dobriban · 2023
Later among the works it cites.
A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit
R Pacelli, S Ariosto, and F Ginelli · 2023
Later among the works it cites.
The rl perceptron: Dynamics of policy learning in high dimensions
Nishil Patel, Sebastian Lee, Stefano Sarao Mannelli, Sebastian Goldt, and Andrew M Saxe · 2023
Later among the works it cites.
Are gaussian data all you need? the extents and limits of universality in high-dimensional generalized linear estimation
Luca Pesce, Florent Krzakala, Bruno Loureiro, and Ludovic Stephan · 2023
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity
Maria Refinetti, Alessandro Ingrosso, and Sebastian Goldt · 2023
Later among the works it cites.
Optimal inference of a generalised Potts model by single-layer transformers with factored attention
Riccardo Rende, Federica Gerace, Alessandro Laio, and Sebastian Goldt · 2023
Later among the works it cites.
Deterministic equivalent and error universality of deep random features learning
Dominik Schröder, Hugo Cui, Daniil Dmitriev, and Bruno Loureiro · 2023
Later among the works it cites.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2023
Later among the works it cites.
Learning curves for deep structured gaussian feature models
Jacob Zavatone-Veth and Cengiz Pehlevan · 2023
Later among the works it cites.
Classification of heavy-tailed features in high dimensions: a superstatistical approach
Urte Adomaityte, Gabriele Sicuro, and Pierpaolo Vivo · 2024
Closest in time.
Random features and polynomial rules
Fabián Aguirre-López, Silvio Franz, and Mauro Pastore · 2024
Closest in time.
Infinite limits of multi-head transformer dynamics
Blake Bordelon, Hamza Tahir Chaudhry, and Cengiz Pehlevan · 2024
Closest in time.
Dynamics of finite width kernel and prediction fluctuations in mean field neural networks
Blake Bordelon and Cengiz Pehlevan · 2024
Closest in time.
The high line: Exact risk and learning rate curves of stochastic adaptive learning rate algorithms
Elizabeth Collins-Woodfin, Inbar Seroussi, Begoña García Malaxechebarría, Andrew W Mackenzie, Elliot Paquette, and Courtney Paquette · 2024
Closest in time.
Hugo Cui, Freya Behrens, Florent Krzakala, and Lenka Zdeborová · 2024
Closest in time.
Analysis of learning a flow-based generative model from limited sample complexity
Hugo Cui, Florent Krzakala, Eric Vanden-Eijnden, and Lenka Zdeborová · 2024
Closest in time.
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala, Yue M Lu, Lenka Zdeborová, and Bruno Loureiro · 2024
Closest in time.
High-dimensional asymptotics of denoising autoencoders
Hugo Cui and Lenka Zdeborová · 2024
Closest in time.
Topics in statistical physics of high-dimensional machine learning
Hugo Chao Cui · 2024
Closest in time.
Universality laws for gaussian mixtures in generalized linear models
Yatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zdeborová · 2024
Closest in time.
Dimension-free deterministic equivalents for random feature regression
Leonardo Defilippis, Bruno Loureiro, and Theodor Misiakiewicz · 2024
Closest in time.
Gaussian universality of perceptrons with random labels
Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborová · 2024
Closest in time.
Rigorous dynamical mean-field theory for stochastic gradient descent methods
Cedric Gerbelot, Emanuele Troiani, Francesca Mignacco, Florent Krzakala, and Lenka Zdeborova · 2024
Closest in time.
Bayesian inference with deep weakly nonlinear networks
Boris Hanin and Alexander Zlokapa · 2024
Closest in time.
Asymptotics of random feature regression beyond the linear scaling regime
Hong Hu, Yue M Lu, and Theodor Misiakiewicz · 2024
Closest in time.
Statistical mechanics of min-max problems
Yuma Ichikawa and Koji Hukushima · 2024
Closest in time.
Asymptotic theory of in-context learning by linear attention
Yue M Lu, Mary I Letey, Jacob A Zavatone-Veth, Anindita Maiti, and Cengiz Pehlevan · 2024
Closest in time.
Bayes-optimal learning of an extensive-width neural network from quadratically many samples
Antoine Maillard, Emanuele Troiani, Simon Martin, Florent Krzakala, and Lenka Zdeborová · 2024
Closest in time.
Asymptotics of learning with deep structured (random) features
Dominik Schröder, Daniil Dmitriev, Hugo Cui, and Bruno Loureiro · 2024
Closest in time.
Dissecting the interplay of attention paths in a statistical mechanics theory of transformers
Lorenzo Tiberi, Francesca Mignacco, Kazuki Irie, and Haim Sompolinsky · 2024
Closest in time.
Coding schemes in neural networks learning classification tasks
Alexander van Meegen and Haim Sompolinsky · 2024
Closest in time.