Fetching the paper…
Reading the bibliography…
We discuss two solvable grokking (generalisation beyond overfitting) models in a rule learning scenario.
Statistical mechanics of cellular automata
Stephen Wolfram · 1983
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
Elizabeth Gardner and Bernard Derrida · 1989
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
A new kind of science
Stephen Wolfram et al · 2002
Earlier work this paper cites.
Supervised learning with quantum-inspired tensor networks
E Miles Stoudenmire and David J Schwab · 2016
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Earlier work this paper cites.
A survey of transfer learning
Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang · 2016
Earlier work this paper cites.
Vasily Pestun and Yiannis Vlassopoulos · 2017
Earlier work this paper cites.
Quantum entanglement in neural network states
Dong-Ling Deng, Xiaopeng Li, and S Das Sarma · 2017
Earlier work this paper cites.
Deep learning and quantum entanglement: Fundamental connections with implications to network design
Yoav Levine, David Yakira, Nadav Cohen, and Amnon Shashua · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Learning relevant features of data with multi-scale tensor networks
E Miles Stoudenmire · 2018
Earlier work this paper cites.
Matrix product operators for sequence-to-sequence learning
Chu Guo, Zhanming Jie, Wei Lu, and Dario Poletti · 2018
Earlier work this paper cites.
Equivalence of restricted boltzmann machines and tensor network states
Jing Chen, Song Cheng, Haidong Xie, Lei Wang, and Tao Xiang · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Tensornetwork for machine learning
Stavros Efthymiou, Jack Hidary, and Stefan Leichenauer · 2019
Earlier work this paper cites.
Machine learning by unitary tensor network of hierarchical tree structure
Ding Liu, Shi-Ju Ran, Peter Wittek, Cheng Peng, Raul Blázquez García, Gang Su, and Maciej Lewenstein · 2019
Earlier work this paper cites.
Tree tensor networks for generative modeling
Song Cheng, Lei Wang, Tao Xiang, and Pan Zhang · 2019
Cited alongside, same era.
Probabilistic modeling with matrix product states
James Stokes and John Terilla · 2019
Cited alongside, same era.
Expressive power of tensor-network factorizations for probabilistic modeling
Ivan Glasser, Ryan Sweke, Nicola Pancotti, Jens Eisert, and Ignacio Cirac · 2019
Cited alongside, same era.
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
Cited alongside, same era.
Machine learning and the physical sciences
Giuseppe Carleo, Ignacio Cirac, Kyle Cranmer, Laurent Daudet, Maria Schuld, Naftali Tishby, Leslie Vogt-Maranto, and Lenka Zdeborová · 2019
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Later among the works it cites.
Yiwei Chen, Yu Pan, and Daoyi Dong · 2021
Later among the works it cites.
Quantum tensor network in machine learning: An application to tiny object classification
Fanjie Kong, Xiao-yang Liu, and Ricardo Henao · 2021
Later among the works it cites.
Tensor networks for unsupervised machine learning
Jing Liu, Sujie Li, Jiang Zhang, and Pan Zhang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vardan Papyan, XY Han, and David L Donoho · 2020
Cited alongside, same era.
Triple descent and the two kinds of overfitting: Where & why do they appear?
Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli · 2020
Cited alongside, same era.
Neural collapse with unconstrained features
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Cited alongside, same era.
Entanglement and tensor networks for supervised image classification
John Martyn, Guifre Vidal, Chase Roberts, and Stefan Leichenauer · 2020
Cited alongside, same era.
Residual matrix product state for machine learning
Ye-Ming Meng, Jing Zhang, Peng Zhang, Chao Gao, and Shi-Ju Ran · 2020
Cited alongside, same era.
Generative tensor network classification model for supervised machine learning
Zheng-Zhi Sun, Cheng Peng, Ding Liu, Shi-Ju Ran, and Gang Su · 2020
Cited alongside, same era.
Modeling sequences with quantum states: a look under the hood
Tai-Danae Bradley, E Miles Stoudenmire, and John Terilla · 2020
Cited alongside, same era.
Later among the works it cites.
Tensor network to learn the wavefunction of data
Anatoly Dymarsky and Kirill Pavlenko · 2021
Later among the works it cites.
Quantum tensor networks, stochastic processes, and weighted automata
Sandesh Adhikary, Siddarth Srinivasan, Jacob Miller, Guillaume Rabusseau, and Byron Boots · 2021
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
Neural collapse: A review on modelling principles and generalization
Vignesh Kothapalli, Ebrahim Rasromani, and Vasudev Awatramani · 2022
Closest in time.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J Michaud, Max Tegmark, and Mike Williams · 2022
Closest in time.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Closest in time.
Extracting finite automata from rnns using state merging
William Merrill and Nikolaos Tsilivis · 2022
Closest in time.
Multi-scale feature learning dynamics: Insights for double descent
Mohammad Pezeshki, Amartya Mitra, Yoshua Bengio, and Guillaume Lajoie · 2022
Closest in time.
Limitations of neural collapse for understanding generalization in deep learning
Like Hui, Mikhail Belkin, and Preetum Nakkiran · 2022
Closest in time.
Deep tensor networks with matrix product operators
Bojan Žunkovič · 2022
Closest in time.
From tensor network quantum states to tensorial recurrent neural networks
Dian Wu, Riccardo Rossi, Filippo Vicentini, and Giuseppe Carleo · 2022
Closest in time.
Wikipedia · 2022
Closest in time.