Fetching the paper…
Reading the bibliography…
In recent years, self-attention has become the dominant paradigm for sequence modeling in a variety of domains.
KiloGrams: Very Large N-Grams for Malware Classification
Raff, E., Fleming, W., Zak, R., Anderson, H., Finlayson, B., Nicholas, C. K., Mclean, M., Fleming, W., Nicholas, C. K., Zak, R., and Mclean, M · 1908
Earlier work this paper cites.
A New Burrows Wheeler Transform Markov Distance
Raff, E., Nicholas, C., and McLean, M · 1912
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P · 1990
Earlier work this paper cites.
Holographic Recurrent Networks
Plate, T. A · 1992
Earlier work this paper cites.
Biologically Inspired Defenses Against Computer Viruses
Kephart, J. O., Sorkin, G. B., Arnold, W. C., Chess, D. M., Tesauro, G. J., and White, S. R · 1995
Earlier work this paper cites.
N-gram-based detection of new malicious code
Abou-Assaleh, T., Cercone, N., Keselj, V., and Sweidan, R · 2004
Earlier work this paper cites.
The Similarity Metric
Li, M., Chen, X., Li, X., Ma, B., and Vitanyi, P. M · 2004
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D., Zhao, Z., and Zheng, C · 2005
Earlier work this paper cites.
Learning to Detect and Classify Malicious Executables in the Wild
Kolter, J. Z. and Maloof, M. A · 2006
Earlier work this paper cites.
Higher-Dimensional Neurons Explain the Tuning and Dynamics of Working Memory Cells
Singh, R. and Eliasmith, C · 2006
Earlier work this paper cites.
Representing word meaning and order information in a composite holographic lexicon
Jones, M. N. and Mewhort, D. J · 2007
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
The Software Similarity Problem in Malware Analysis
Walenstein, A. and Lakhotia, A · 2007
Earlier work this paper cites.
Multi-class pegasos on a budget
Wang, Z., Crammer, K., and Vucetic, S · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Malware Images: Visualization and Automatic Classification
Nataraj, L., Karthikeyan, S., Jacob, G., and Manjunath, B. S · 2011
Earlier work this paper cites.
Training Transformers for Information Security Tasks: A Case Study on Malicious URL Prediction
Rudd, E. M. and Abdallah, A · 2011
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2011
Earlier work this paper cites.
Trading representability for scalability Adaptive Multi-Hyperplane Machine for nonlinear Classification
Wang, Z., Djuric, N., Crammer, K., and Vucetic, S · 2011
Earlier work this paper cites.
A Large-Scale Model of the Functioning Brain
Eliasmith, C., Stewart, T. C., Choo, X., Bekolay, T., DeWolf, T., Tang, Y., and Rasmussen, D · 2012
Earlier work this paper cites.
Classifying Sequences of Extreme Length with Constant Memory Applied to Malware Detection
Raff, E., Fleshman, W., Zak, R., Anderson, H. S., Filar, B., and McLean, M · 2012
Earlier work this paper cites.
Nengo: a Python tool for building large-scale functional brain models
Bekolay, T., Bergstra, J., Hunsberger, E., DeWolf, T., Stewart, T., Rasmussen, D., Choo, X., Voelker, A., and Eliasmith, C · 2013
Earlier work this paper cites.
A Neurally Plausible Encoding of Word Order Information into a Semantic Vector Space
Blouw, P. and Eliasmith, C · 2013
Earlier work this paper cites.
The acl anthology network corpus
Radev, D. R., Muthukrishnan, P., Qazvinian, V., and Abu-Jbara, A · 2013
Cited alongside, same era.
Large-margin Convex Polytope Machine
Kantchelian, A., Tschantz, M. C., Huang, L., Bartlett, P. L., Joseph, A. D., and Tygar, J. D · 2014
Cited alongside, same era.
Large-scale synthesis of functional spiking neural circuits
Stewart, T. C. and Eliasmith, C · 2014
Cited alongside, same era.
Scaling SVM and Least Absolute Deviations via Exact Data Reduction
Wang, J., Wonka, P., and Ye, J · 2014
Cited alongside, same era.
On normalized compression distance and large malware
Borbely, R. S · 2015
Cited alongside, same era.
Improved Bounds on the Dot Product under Random Projection and Random Sign Projection
Kaban, A · 2015
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Later among the works it cites.
A Survey on Using Kolmogorov Complexity in Cybersecurity
S. Resende, J., Martins, R., and Antunes, L · 2019
Later among the works it cites.
Adaptive Attention Span in Transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks
Voelker, A., Kajić, I., and Eliasmith, C · 2019
Later among the works it cites.
When Malware is Packin’ Heat; Limits of Machine Learning Classifiers Based on Static Analysis Features
Aghakhani, H., Gritti, F., Mecca, F., Lindorfer, M., Ortolani, S., Balzarotti, D., Vigna, G., and Kruegel, C · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Concepts as Semantic Pointers: A Framework and Computational Model
Blouw, P., Solodkin, E., Thagard, P., and Eliasmith, C · 2016
Cited alongside, same era.
Associative Long Short-Term Memory
Danihelka, I., Wayne, G., Uria, B., Kalchbrenner, N., and Graves, A · 2016
Cited alongside, same era.
Malware classification using gray-scale images and ensemble learning
Liu, L. and Wang, B · 2016
Cited alongside, same era.
Holographic Embeddings of Knowledge Graphs
Nickel, M., Rosasco, L., and Poggio, T · 2016
Cited alongside, same era.
Computationally Efficient Nystrom Approximation using Fast Transforms
Si, S., Hsieh, C.-J., and Dhillon, I. S · 2016
Cited alongside, same era.
Learning Kernels with Random Features
Sinha, A. and Duchi, J. C · 2016
Cited alongside, same era.
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Later among the works it cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Later among the works it cites.
HiPPO: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Later among the works it cites.
Self-Attentive Associative Memory
Le, H., Tran, T., and Venkatesh, S · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
Learning with Holographic Reduced Representations
Ganesan, A., Gao, H., Gandhi, S., Raff, E., Oates, T., Holt, J., and McLean, M · 2021
Later among the works it cites.
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., and Ré, C · 2021
Later among the works it cites.
FNet: Mixing Tokens with Fourier Transforms
Lee-Thorp, J., Ainslie, J., Eckstein, I., and Ontanon, S · 2021
Later among the works it cites.
Luna: Linear Unified Nested Attention
Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L · 2021
Later among the works it cites.
Nystr\”omformer: A Nystr\”om-Based Algorithm for Approximating Self-Attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Later among the works it cites.
H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
Zhu, Z. and Soricut, R · 2021
Later among the works it cites.
It’s Raw! Audio Generation with State-Space Models
Goel, K., Gu, A., Donahue, C., and Ré, C · 2022
Later among the works it cites.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A., Goel, K., and Ré, C · 2022
Later among the works it cites.
Transformers for End-to-End InfoSec Tasks: A Feasibility Study
Rudd, E. M., Rahman, M. S., and Tully, P · 2022
Later among the works it cites.