Fetching the paper…
Reading the bibliography…
Transformers have emerged as a preferred model for many tasks in natural langugage processing and vision.
Multi-Grid Methods and Applications , volume 4
Hackbusch, W · 1985
Earlier work this paper cites.
Ten Lectures on Wavelets
Daubechies, I · 1992
Earlier work this paper cites.
De-noising by soft-thresholding
Donoho, D. L · 1995
Earlier work this paper cites.
Adapting to unknown smoothness via wavelet shrinkage
Donoho, D. L. and Johnstone, I. M · 1995
Earlier work this paper cites.
A Wavelet Tour of Signal Processing
Mallat, S · 1999
Earlier work this paper cites.
Iterative Methods for Sparse Linear Systems
Saad, Y · 2003
Earlier work this paper cites.
Diffusion wavelets
Coifman, R. R. and Maggioni, M · 2006
Earlier work this paper cites.
Treelets: A tool for dimensionality reduction and multi-scale analysis of unstructured data
Lee, A. B. and Nadler, B · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Jensen’s inequality and new entropy bounds
Simic, S · 2009
Earlier work this paper cites.
Multiscale wavelets on trees, graphs and high dimensional data: Theory and applications to semi supervised learning
Gavish, M., Nadler, B., and Coifman, R. R · 2010
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
Wavelets on graphs via spectral graph theory
Hammond, D. K., Vandergheynst, P., and Gribonval, R · 2011
Earlier work this paper cites.
Multiresolution matrix factorization
Kondor, R., Teneva, N., and Garg, V · 2014
Earlier work this paper cites.
Relations of the nuclear norm of a tensor and its matrix flattenings
Hu, S · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
The incremental multiresolution matrix factorization algorithm
Ithapu, V. K., Kondor, R., Johnson, S. C., and Singh, V · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Connections between nuclear-norm and frobenius-norm-based representations
Peng, X., Lu, C., Yi, Z., and Tang, H · 2018
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Later among the works it cites.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Later among the works it cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trinh, T. H. and Le, Q. V · 2018
Cited alongside, same era.
Constructing datasets for multi-hop reading comprehension across documents
Welbl, J., Stenetorp, P., and Riedel, S · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Revealing the dark secrets of bert
Kovaleva, O., Romanov, A., Rogers, A., and Rumshisky, A · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Defending against neural fake news
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y · 2019
Cited alongside, same era.
Convit: Improving vision transformers with soft convolutional inductive biases
d’Ascoli, S., Touvron, H., Leavitt, M., Morcos, A., Biroli, G., and Sagun, L · 2021
Later among the works it cites.
Fnet: Mixing tokens with fourier transforms
Lee-Thorp, J., Ainslie, J., Eckstein, I., and Ontanon, S · 2021
Later among the works it cites.
Soft: Softmax-free transformer with linear complexity
Lu, J., Yao, J., Zhang, J., Zhu, X., Xu, H., Gao, W., Xu, C., Xiang, T., and Zhang, L · 2021
Later among the works it cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Later among the works it cites.
Long range arena : A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Later among the works it cites.
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Xu, Y., Zhang, Q., Zhang, J., and Tao, D · 2021
Later among the works it cites.
You only sample (almost) once: Linear cost self-attention via bernoulli sampling
Zeng, Z., Xiong, Y., Ravi, S., Acharya, S., Fung, G. M., and Singh, V · 2021
Later among the works it cites.
H-transformer-1D: Fast one-dimensional hierarchical attention for sequences
Zhu, Z. and Soricut, R · 2021
Later among the works it cites.