Fetching the paper…
Reading the bibliography…
We present a nonparametric interpretation for deep learning compatible modern Hopfield models and utilize this new perspective to debut efficient variants.
Characteristics of random nets of analog neuron-like elements
Amari, S.-I · 1972
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J · 1982
Earlier work this paper cites.
Neurons with graded response have collective computational properties like those of two-state neurons
Hopfield, J. J · 1984
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2002
Earlier work this paper cites.
On the convergence of the concave-convex procedure
Sriperumbudur, B. K. and Lanckriet, G. R · 2009
Earlier work this paper cites.
NIST handbook of mathematical functions hardback and CD-ROM
Olver, F. W., Lozier, D. W., Boisvert, R. F., and Clark, C. W · 2010
Earlier work this paper cites.
Random feature maps for dot product kernels
Kar, P. and Karnick, H · 2012
Earlier work this paper cites.
The nature of statistical learning theory
Vapnik, V · 2013
Earlier work this paper cites.
Compact random feature maps
Hamid, R., Xiao, Y., Gittens, A., and DeCoste, D · 2014
Earlier work this paper cites.
An equivalence between the lasso and support vector machines
Jaggi, M · 2014
Earlier work this paper cites.
Empowering multiple instance histopathology cancer diagnosis by cell graphs
Kandemir, M., Zhang, C., and Hamprecht, F. A · 2014
Earlier work this paper cites.
Support vector regression
Awad, M., Khanna, R., Awad, M., and Khanna, R · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Dense associative memory for pattern recognition
Krotov, D. and Hopfield, J. J · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, A. and Astudillo, R · 2016
Earlier work this paper cites.
On a model of associative memory with huge storage capacity
Demircigil, M., Heusel, J., Löwe, M., Upgang, S., and Vermet, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Random point sets on the sphere-hole radii, covering, and separation
Brauchart, J. S., Reznikov, A. B., Saff, E. B., Sloan, I. H., Wang, Y. G., and Womersley, R. S · 2018
Earlier work this paper cites.
Multiple instance learning: A survey of problem characteristics and applications
Carbonneau, M.-A., Cheplygina, V., Granger, E., and Gagnon, G · 2018
Earlier work this paper cites.
Attention-based deep multiple instance learning
Ilse, M., Tomczak, J., and Welling, M · 2018
Cited alongside, same era.
Sparsemap: Differentiable sparse structured inference
Niculae, V., Martins, A., Blondel, M., and Cardie, C · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Adaptively sparse transformers
Correia, G. M., Niculae, V., and Martins, A. F · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, S. W.-t., Wang, S., and Tang, J · 2019
Cited alongside, same era.
Cloob: Modern hopfield networks with infoloob outperform clip
Fürst, A., Rumetshofer, E., Lehner, J., Tran, V. T., Tang, F., Ramsauer, H., Kreil, D., Kopp, M., Klambauer, G., Bitto, A., et al · 2022
Later among the works it cites.
Kernel memory networks: A unifying framework for memory modeling
Iatropoulos, G., Brea, J., and Gerstner, W · 2022
Later among the works it cites.
Building transformers from neurons and astrocytes
Kozachkov, L., Kastanenka, K. V., and Krotov, D · 2022
Later among the works it cites.
History compression via language models in reinforcement learning
Paischer, F., Adler, T., Patil, V., Bitto-Nemling, A., Holzleitner, M., Lehner, S., Eghbal-Zadeh, H., and Hochreiter, S · 2022
Later among the works it cites.
Improving few-and zero-shot reaction template prediction using modern hopfield networks
Seidl, P., Renz, P., Dyubankova, N., Neves, P., Verhoeven, J., Wegner, J. K., Segler, M., Hochreiter, S., and Klambauer, G · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Gpt-3: Its nature, scope, limits, and consequences
Floridi, L. and Chiriatti, M · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Hopfield networks is all you need
Ramsauer, H., Schafl, B., Lehner, J., Seidl, P., Widrich, M., Adler, T., Gruber, L., Holzleitner, M., Pavlovic, M., Sandve, G. K., et al · 2020
Cited alongside, same era.
Modern hopfield networks and attention for immune repertoire classification
Widrich, M., Schäfl, B., Pavlović, M., Ramsauer, H., Gruber, L., Holzleitner, M., Brandstetter, J., Sandve, G. K., Greiff, V., Hochreiter, S., et al · 2020
Cited alongside, same era.
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2022
Later among the works it cites.
Blog post: Hopfield networks is all you need, 2021
Brandstetter, J · 2023
Later among the works it cites.
Superiority of softmax: Unveiling the performance edge over linear attention
Deng, Y., Song, Z., and Zhou, T · 2023
Later among the works it cites.
Hoover, B., Liang, Y., Pham, B., Panda, R., Strobelt, H., Chau, D. H., Zaki, M. J., and Krotov, D · 2023
Later among the works it cites.
On sparse modern hopfield model
Hu, J. Y.-C., Yang, D., Wu, D., Xu, C., Chen, B.-Y., and Liu, H · 2023
Later among the works it cites.
On the computational complexity of self-attention
Keles, F. D., Wijewardena, P. M., and Hegde, C · 2023
Later among the works it cites.
Sparse modern hopfield networks
Martins, A. F. T., Niculae, V., and McNamee, D · 2023
Later among the works it cites.
Context-enriched molecule representations improve few-shot drug discovery
Schimunek, J., Seidl, P., Friedrich, L., Kuhn, D., Rippmann, F., Hochreiter, S., and Klambauer, G · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., and Mann, G · 2023
Later among the works it cites.
Conformal prediction for time series with modern hopfield networks
Auer, A., Gauch, M., Klotz, D., and Hochreiter, S · 2024
Closest in time.
Dense associative memory through the lens of random features
Hoover, B., Chau, D. H., Strobelt, H., Ram, P., and Krotov, D · 2024
Closest in time.
A primal-dual framework for transformers and neural networks
Nguyen, T. M., Nguyen, T., Ho, N., Bertozzi, A. L., Baraniuk, R. G., and Osher, S. J · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Xu, C., Huang, Y.-C., Hu, J. Y.-C., Li, W., Gilani, A., Goan, H.-S., and Liu, H · 2024
Closest in time.