Fetching the paper…
Reading the bibliography…
We introduce the sparse modern Hopfield model as a sparse extension of the modern Hopfield model.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André FT Martins · 1905
Earlier work this paper cites.
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 1909
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 1909
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang · 1911
Earlier work this paper cites.
Nonlinear programming: a unified approach , volume 52
Willard I Zangwill · 1969
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Neurons with graded response have collective computational properties like those of two-state neurons
John J Hopfield · 1984
Earlier work this paper cites.
Machine learning using a higher order correlation network
YC Lee, Gary Doolen, HH Chen, GZ Sun, Tom Maxwell, and HY Lee · 1986
Earlier work this paper cites.
Long term memory storage capacity of multiconnected neural networks
Pierre Peretto and Jean-Jacques Niez · 1986
Earlier work this paper cites.
Memory capacity in neural network models: Rigorous lower bounds
Charles M Newman · 1988
Earlier work this paper cites.
Forming sparse representations by local anti-hebbian learning
Peter Földiak · 1990
Earlier work this paper cites.
On the lambert w function
Robert M Corless, Gaston H Gonnet, David EG Hare, David J Jeffrey, and Donald E Knuth · 1996
Earlier work this paper cites.
Solving the multiple instance problem with axis-parallel rectangles
Thomas G Dietterich, Richard H Lathrop, and Tomás Lozano-Pérez · 1997
Earlier work this paper cites.
A framework for multiple-instance learning
Oded Maron and Tomás Lozano-Pérez · 1997
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Solving multiple-instance problem: A lazy learning approach
Jun Wang and Jean-Daniel Zucker · 2000
Earlier work this paper cites.
The concave-convex procedure (cccp)
Alan L Yuille and Anand Rangarajan · 2001
Earlier work this paper cites.
The concave-convex procedure
Alan L Yuille and Anand Rangarajan · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2004
Earlier work this paper cites.
Order statistics
Herbert A David and Haikady N Nagaraja · 2004
Earlier work this paper cites.
Convergence theorems for generalized alternating minimization procedures
Asela Gunawardana, William Byrne, and Michael I Jordan · 2005
Earlier work this paper cites.
Miles: Multiple-instance learning via embedded instance selection
Yixin Chen, Jinbo Bi, and James Ze Wang · 2006
Earlier work this paper cites.
Modern hopfield networks and attention for immune repertoire classification
Michael Widrich, Bernhard Schäfl, Milena Pavlović, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Johannes Brandstetter, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, et al · 2007
Cited alongside, same era.
Large associative memory problem in neurobiology and machine learning
Dmitry Krotov and John J. Hopfield · 2008
Cited alongside, same era.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David Kreil, Michael K Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2008
Cited alongside, same era.
Graphical models, exponential families, and variational inference
Martin J Wainwright, Michael I Jordan, et al · 2008
Cited alongside, same era.
On the convergence of the concave-convex procedure
Random point sets on the sphere—hole radii, covering, and separation
Johann S Brauchart, Alexander B Reznikov, Edward B Saff, Ian H Sloan, Yu Guang Wang, and Robert S Womersley · 2018
Later among the works it cites.
Multiple instance learning: A survey of problem characteristics and applications
Marc-André Carbonneau, Veronika Cheplygina, Eric Granger, and Ghyslain Gagnon · 2018
Later among the works it cites.
Attention-based deep multiple instance learning
Maximilian Ilse, Jakub Tomczak, and Max Welling · 2018
Later among the works it cites.
Bag encoding strategies in multiple instance learning problems
Emel Şeyma Küçükaşcı and Mustafa Gökçe Baydoğan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bharath K Sriperumbudur and Gert RG Lanckriet · 2009
Cited alongside, same era.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2009
Cited alongside, same era.
Sparse and redundant representations: from theory to applications in signal and image processing , volume 2
Michael Elad · 2010
Cited alongside, same era.
Online learning for matrix factorization and sparse coding
Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro · 2010
Cited alongside, same era.
NIST handbook of mathematical functions hardback and CD-ROM
Frank WJ Olver, Daniel W Lozier, Ronald F Boisvert, and Charles W Clark · 2010
Cited alongside, same era.
Dictionaries for sparse representation modeling
Ron Rubinstein, Alfred M Bruckstein, and Michael Elad · 2010
Cited alongside, same era.
Phase transition in limiting distributions of coherence of high-dimensional random matrices
T Tony Cai and Tiefeng Jiang · 2012
Cited alongside, same era.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2012
Cited alongside, same era.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome
Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V Davuluri · 2020
Later among the works it cites.
Sparse sinkhorn attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2020
Later among the works it cites.
Autoformer: Searching transformers for visual recognition
Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Cloob: Modern hopfield networks with infoloob outperform clip
Andreas Fürst, Elisabeth Rumetshofer, Johannes Lehner, Viet T Tran, Fei Tang, Hubert Ramsauer, David Kreil, Michael Kopp, Günter Klambauer, Angela Bitto, et al · 2022
Later among the works it cites.
Building transformers from neurons and astrocytes
Leo Kozachkov, Ksenia V Kastanenka, and Dmitry Krotov · 2022
Later among the works it cites.
Universal hopfield networks: A general framework for single-shot associative memory models
Beren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz, and Rafal Bogacz · 2022
Later among the works it cites.
History compression via language models in reinforcement learning
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbal-Zadeh, and Sepp Hochreiter · 2022
Later among the works it cites.
Improving few-and zero-shot reaction template prediction using modern hopfield networks
Philipp Seidl, Philipp Renz, Natalia Dyubankova, Paulo Neves, Jonas Verhoeven, Jorg K Wegner, Marwin Segler, Sepp Hochreiter, and Gunter Klambauer · 2022
Later among the works it cites.
Transformers from an optimization perspective
Yongyi Yang, Zengfeng Huang, and David Wipf · 2022
Later among the works it cites.
Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting
Yunhao Zhang and Junchi Yan · 2022
Later among the works it cites.
Film: Frequency improved legendre memory model for long-term time series forecasting
Tian Zhou, Ziqing Ma, Qingsong Wen, Liang Sun, Tao Yao, Wotao Yin, Rong Jin, et al · 2022
Later among the works it cites.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Anonymous · 2023
Closest in time.
Blog post: Hopfield networks is all you need, 2021
Johannes Brandstetter · 2023
Closest in time.
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed J Zaki, and Dmitry Krotov · 2023
Closest in time.
Sparse modern hopfield networks
Andre F. T. Martins, Vlad Niculae, and Daniel McNamee · 2023
Closest in time.