Fetching the paper…
Reading the bibliography…
We propose a new class of linear Transformers called FourierLearner-Transformers (FLTs), which incorporate a wide range of relative positional encoding mechanisms (RPEs).
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B. (2007) · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning kernels with random features
Sinha, A. and Duchi, J. C. (2016) · 2016
Earlier work this paper cites.
Random fourier features for kernel ridge regression: Approximation bounds and statistical guarantees
Avron, H., Kapralov, M., Musco, C., Musco, C., Velingker, A., and Zandieh, A. (2017) · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R. (2017) · 2017
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A. (2018) · 2018
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, À., Khosla, A., Oliva, A., and Torralba, A. (2018) · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R. (2019) · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. (2019) · 2019
Earlier work this paper cites.
On kernel derivative approximation with random fourier features
Szabó, Z. and Sriperumbudur, B. K. (2019) · 2019
Earlier work this paper cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
Open catalyst 2020 (oc20) dataset and community challenges
Chanussot*, L., Das*, A., Goyal*, S., Lavril*, T., Shuaibi*, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. (2021) · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. (2020) · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A. (2020) · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Earlier work this paper cites.
Language through a prism: A spectral approach for multiscale language representations
Tamkin, A., Jurafsky, D., and Goodman, N. (2020) · 2020
Earlier work this paper cites.
Fast transformers with clustered attention
Vyas, A., Katharopoulos, A., and Fleuret, F. (2020) · 2020
Earlier work this paper cites.
O(n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S. (2020) · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al. (2020) · 2020
Cited alongside, same era.
Lambdanetworks: Modeling long-range interactions without attention
Bello, I. (2021) · 2021
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Cited alongside, same era.
Gemnet: Universal directional graph neural networks for molecules
Gasteiger, J., Becker, F., and Günnemann, S. (2021) · 2021
Cited alongside, same era.
Effective gene expression prediction from sequence by integrating long-range interactions
Žiga Avsec, Agarwal, V., Visentin, D., Ledsam, J. R., Grabska-Barwinska, A., Taylor, K. R., Assael, Y., Jumper, J. M., Kohli, P., and Kelley, D. R. (2021) · 2021
Later among the works it cites.
From block-Toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked transformers
Choromanski, K., Lin, H., Chen, H., Zhang, T., Sehanobish, A., Likhosherstov, V., Parker-Holder, J., Sarlós, T., Weller, A., and Weingarten, T. (2022a) · 2022
Later among the works it cites.
Hybrid random features
Choromanski, K. M., Lin, H., Chen, H., Sehanobish, A., Ma, Y., Jain, D., Varley, J., Zeng, A., Ryoo, M. S., Likhosherstov, V., Kalashnikov, D., Sindhwani, V., and Weller, A. (2022b) · 2022
Later among the works it cites.
On learning the transformer kernel
Chowdhury, S. P., Solomou, A., Dubey, A., and Sachan, M. (2022) · 2022
Later among the works it cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geng, Z., Guo, M.-H., Chen, H., Li, X., Wei, K., and Lin, Z. (2021) · 2021
Cited alongside, same era.
Translational equivariance in kernelizable attention
Horn, M., Shridhar, K., Groenewald, E., and Baumann, P. F. M. (2021) · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. (2021) · 2021
Cited alongside, same era.
Can vision transformers perform convolution?
Li, S., Chen, X., He, D., and Hsieh, C.-J. (2021) · 2021
Cited alongside, same era.
Quantization algorithms for random fourier features
Li, X. and Li, P. (2021) · 2021
Cited alongside, same era.
Relative positional encoding for transformers with linear complexity
Liutkus, A., Cífka, O., Wu, S., Simsekli, U., Yang, Y., and Richard, G. (2021) · 2021
Cited alongside, same era.
Stable, fast and accurate: Kernelized attention with relative positional encoding
Luo, S., Li, S., Cai, T., He, D., Peng, D., Zheng, S., Ke, G., Wang, L., and Liu, T.-Y. (2021) · 2021
Cited alongside, same era.
Chefs'random tables: Non-trigonometric random features
Likhosherstov, V., Choromanski, K. M., Dubey, K. A., Liu, F., Sarlos, T., and Weller, A. (2022) · 2022
Later among the works it cites.
Your transformer may not be as powerful as you expect
Luo, S., Li, S., Zheng, S., Liu, T.-Y., Wang, L., and He, D. (2022) · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M. (2022) · 2022
Later among the works it cites.
cosformer: Rethinking softmax in attention
Qin, Z., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y. (2022) · 2022
Later among the works it cites.
Benchmarking graphormer on large-scale molecular modeling datasets
Shi, Y., Zheng, S., Ke, G., Shen, Y., You, J., He, J., Luo, S., Liu, C., He, D., and Liu, T.-Y. (2022) · 2022
Later among the works it cites.
Sparse attention with learning to hash
Sun, Z., Yang, Y., and Yoo, S. (2022) · 2022
Later among the works it cites.
Learning model predictive controllers with real-time attention for real-world navigation
Xiao, X., Zhang, T., Choromanski, K., Lee, T. E., Francis, A. G., Varley, J., Tu, S., Singh, S., Xu, P., Xia, F., Persson, S. M., Kalashnikov, D., Takayama, L., Frostig, R., Tan, J., Parada, C., and Sindhwani, V. (2022) · 2022
Later among the works it cites.
Transformer-based learned optimization
Gärtner, E., Metz, L., Andriluka, M., Freeman, C. D., and Sminchisescu, C. (2023) · 2023
Closest in time.
Mnemosyne: Learning to train transformers with transformers
Jain, D., Choromanski, K. M., Dubey, K. A., Singh, S., Sindhwani, V., Zhang, T., and Tan, J. (2023) · 2023
Closest in time.
One transformer can understand both 2d & 3d molecular data
Luo, S., Chen, T., Xu, Y., Zheng, S., Liu, T.-Y., Wang, L., and He, D. (2023) · 2023
Closest in time.
Deep autoregressive models with spectral attention
Moreno-Pino, F., Olmos, P. M., and Artés-Rodríguez, A. (2023) · 2023
Closest in time.
Randomized positional encodings boost length generalization of transformers
Ruoss, A., Delétang, G., Genewein, T., Grau-Moya, J., Csordás, R., Bennani, M., Legg, S., and Veness, J. (2023) · 2023
Closest in time.
Functional interpolation for relative positions improves long context transformers
Li, S., You, C., Guruganesh, G., Ainslie, J., Ontanon, S., Zaheer, M., Sanghai, S., Yang, Y., Kumar, S., and Bhojanapalli, S. (2024) · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024) · 2024
Closest in time.
Do efficient transformers really save computation?
Yang, K., Ackermann, J., He, Z., Feng, G., Zhang, B., Feng, Y., Ye, Q., He, D., and Wang, L. (2024) · 2024
Closest in time.