Fetching the paper…
Reading the bibliography…
Modern Hopfield Networks (MHNs) have emerged as powerful components in deep learning, serving as effective replacements for pooling layers, LSTMs, and attention mechanisms.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Neurons with graded response have collective computational properties like those of two-state neurons
John J Hopfield · 1984
Earlier work this paper cites.
A taxonomy of problems with fast parallel algorithms
Stephen A Cook · 1985
Earlier work this paper cites.
Problems complete for deterministic logarithmic space
Stephen A Cook and Pierre McKenzie · 1987
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Descriptive complexity
Neil Immerman · 1998
Earlier work this paper cites.
A note on the hardness of tree isomorphism
Birgit Jenner, Pierre McKenzie, and Jacobo Torán · 1998
Earlier work this paper cites.
Introduction to circuit complexity: a uniform approach
Heribert Vollmer · 1999
Earlier work this paper cites.
On the complexity of k-sat
Russell Impagliazzo and Ramamohan Paturi · 2001
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, 2014
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate, 2016
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Permutation-equivariant neural networks applied to dynamics prediction
Nicholas Guttenberg, Nathaniel Virgo, Olaf Witkowski, Hidetoshi Aoki, and Ryota Kanai · 2016
Earlier work this paper cites.
Dense associative memory for pattern recognition
Dmitry Krotov and John J Hopfield · 2016
Earlier work this paper cites.
Deep learning with sets and point clouds
Siamak Ravanbakhsh, Jeff Schneider, and Barnabas Poczos · 2016
Earlier work this paper cites.
On a model of associative memory with huge storage capacity
Mete Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck Vermet · 2017
Earlier work this paper cites.
Dense associative memory is robust to adversarial inputs
Dmitry Krotov and John Hopfield · 2018
Cited alongside, same era.
Large associative memory problem in neurobiology and machine learning
Dmitry Krotov and John Hopfield · 2020
Cited alongside, same era.
Modern hopfield networks and attention for immune repertoire classification, 2020
Michael Widrich, Bernhard Schäfl, Hubert Ramsauer, Milena Pavlović, Lukas Gruber, Markus Holzleitner, Johannes Brandstetter, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, and Günter Klambauer · 2020
Cited alongside, same era.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schafl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David Kreil, Michael K Kopp, et al · 2021
Cited alongside, same era.
What is my math transformer doing?–three results on interpretability and generalization
François Charton · 2022
Cited alongside, same era.
History compression via language models in reinforcement learning, 2023
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbal-zadeh, and Sepp Hochreiter · 2023
Later among the works it cites.
Context-enriched molecule representations improve few-shot drug discovery
Johannes Schimunek, Philipp Seidl, Lukas Friedrich, Daniel Kuhn, Friedrich Rippmann, Sepp Hochreiter, and Günter Klambauer · 2023
Later among the works it cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Later among the works it cites.
The fine-grained complexity of gradient computation for training large language models
Josh Alman and Zhao Song · 2024
Closest in time.
David Chiang · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cloob: Modern hopfield networks with infoloob outperform clip, 2022
Andreas Fürst, Elisabeth Rumetshofer, Johannes Lehner, Viet Tran, Fei Tang, Hubert Ramsauer, David Kreil, Michael Kopp, Günter Klambauer, Angela Bitto-Nemling, and Sepp Hochreiter · 2022
Cited alongside, same era.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Cited alongside, same era.
Improving few-and zero-shot reaction template prediction using modern hopfield networks
Philipp Seidl, Philipp Renz, Natalia Dyubankova, Paulo Neves, Jonas Verhoeven, Jorg K Wegner, Marwin Segler, Sepp Hochreiter, and Gunter Klambauer · 2022
Cited alongside, same era.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Cited alongside, same era.
Josh Alman and Zhao Song · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Circuit complexity bounds for rope-based transformer architecture
Bo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
Jerry Yao-Chieh Hu, Pei-Hsuan Chang, Haozheng Luo, Hong-Yu Chen, Weijian Li, Wei-Po Wang, and Han Liu · 2024
Closest in time.
Nonparametric modern hopfield models
Jerry Yao-Chieh Hu, Bo-Yu Chen, Dennis Wu, Feng Ruan, and Han Liu · 2024
Closest in time.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Closest in time.
Provably optimal memory capacity for modern hopfield models: Transformer-compatible dense associative memories as spherical codes
Jerry Yao-Chieh Hu, Dennis Wu, and Han Liu · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.