Fetching the paper…
Reading the bibliography…
We introduce an Outlier-Efficient Modern Hopfield Model (termed $\mathrm{OutEffHop}$) and use it to address the outlier inefficiency problem of {training} gigantic transformer-based models.
Efficient 8-bit quantization of transformer neural machine language translation model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore · 1906
Earlier work this paper cites.
Revealing the dark secrets of bert, 2019
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky · 1908
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 1910
Earlier work this paper cites.
Central limit theorems for empirical measures
Richard M Dudley · 1978
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Neurons with graded response have collective computational properties like those of two-state neurons
John J Hopfield · 1984
Earlier work this paper cites.
Fast neural networks without multipliers
M. Marchesi, G. Orlandi, F. Piazza, and A. Uncini · 1993
Earlier work this paper cites.
Multilayer feedforward neural networks with single powers-of-two weights
C.Z. Tang and H.K. Kwan · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Robust quantization: One model to rule them all, 2020
Moran Shkolnik, Brian Chmiel, Ron Banner, Gil Shomron, Yury Nahshan, Alex Bronstein, and Uri Weiser · 2002
Earlier work this paper cites.
The Concave-Convex Procedure
A. L. Yuille and Anand Rangarajan · 2003
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2005
Earlier work this paper cites.
Modern hopfield networks and attention for immune repertoire classification
Michael Widrich, Bernhard Schäfl, Milena Pavlović, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Johannes Brandstetter, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, et al · 2007
Earlier work this paper cites.
Large associative memory problem in neurobiology and machine learning
Dmitry Krotov and John J. Hopfield · 2008
Earlier work this paper cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schafl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlovic, Geir Kjetil Sandve, et al · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
On the convergence of the concave-convex procedure
Bharath K Sriperumbudur and Gert RG Lanckriet · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2010
Earlier work this paper cites.
NIST handbook of mathematical functions hardback and CD-ROM
Frank WJ Olver, Daniel W Lozier, Ronald F Boisvert, and Charles W Clark · 2010
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting, 2021
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2012
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Dense associative memory for pattern recognition
Dmitry Krotov and John J. Hopfield · 2016
Cited alongside, same era.
Cs229t/stat231: Statistical learning theory (winter 2016), 2016
Percy Liang · 2016
Cited alongside, same era.
On a model of associative memory with huge storage capacity
Mete Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck Vermet · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Later among the works it cites.
Conformal prediction for time series with modern hopfield networks
Andreas Auer, Martin Gauch, Daniel Klotz, and Sepp Hochreiter · 2023
Later among the works it cites.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Later among the works it cites.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort · 2023
Later among the works it cites.
Blog post: Hopfield networks is all you need, 2021
Johannes Brandstetter · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Marian: Cost-effective high-quality neural machine translation in c++
Marcin Junczys-Dowmunt, Kenneth Heafield, Hieu Hoang, Roman Grundkiewicz, and Anthony Aue · 2018
Cited alongside, same era.
revealt does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Gpt-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti · 2020
Cited alongside, same era.
Wiki-40b: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou · 2020
Cited alongside, same era.
Attention is not only a weight: Analyzing transformers with vector norms, 2020
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Simplicial hopfield networks
Thomas F Burns and Tomoki Fukai · 2023
Later among the works it cites.
Hamza Chaudhry, Jacob Zavatone-Veth, Dmitry Krotov, and Cengiz Pehlevan · 2023
Later among the works it cites.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Later among the works it cites.
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed J Zaki, and Dmitry Krotov · 2023
Later among the works it cites.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Later among the works it cites.
Blog post: Exploring softmax1, or “community research for the win!”, 2023
johnowhitaker · 2023
Later among the works it cites.
Blog post: Attention is off by one, 2023
Evan Miller · 2023
Later among the works it cites.
Feature programming for multivariate time series prediction
Alex Reneau, Jerry Yao-Chieh Hu, Chenwei Xu, Weijian Li, Ammar Gilani, and Han Liu · 2023
Later among the works it cites.
Context-enriched molecule representations improve few-shot drug discovery
Johannes Schimunek, Philipp Seidl, Lukas Friedrich, Daniel Kuhn, Friedrich Rippmann, Sepp Hochreiter, and Günter Klambauer · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
Mathematical analysis of machine learning algorithms
Tong Zhang · 2023
Later among the works it cites.
Dnabert-2: Efficient foundation model and benchmark for multi-species genome
Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu · 2023
Later among the works it cites.
Semantically-correlated memories in a dense associative model
Thomas F Burns · 2024
Closest in time.
Scaling laws for associative memories
Vivien Cabannes, Elvis Dohmatob, and Alberto Bietti · 2024
Closest in time.
Energy-based hopfield boosting for out-of-distribution detection
Claus Hofmann, Simon Schmid, Bernhard Lehner, Daniel Klotz, and Sepp Hochreiter · 2024
Closest in time.
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Dnabert-s: Learning species-aware dna embedding with genome foundation models
Zhihan Zhou, Weimin Wu, Harrison Ho, Jiayi Wang, Lizhen Shi, Ramana V Davuluri, Zhong Wang, and Han Liu · 2024
Closest in time.