Fetching the paper…
Reading the bibliography…
Transformer-based language models utilize the attention mechanism for substantial performance improvements in almost all natural language processing (NLP) tasks.
T. Brown et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901
1901
Earlier work this paper cites.
A. R. Gallant and H. White, “There exists a neural network that does not make avoidable mistakes,” in IEEE 1988 International Conference on Neural Networks , 1988, pp. 657–664 vol.1
1988
Earlier work this paper cites.
R. Koplon and E. D. Sontag, “Using Fourier-neural recurrent networks to fit sequential input/output data,” Neurocomputing , vol. 15, no. 3, pp. 225–248, 1997
1997
Earlier work this paper cites.
A. Silvescu, “Fourier neural networks,” in IJCNN’99. International Joint Conference on Neural Networks. Proceedings , vol. 1, 1999, pp. 488–491 vol.1
1999
Earlier work this paper cites.
S. Ben-Yacoub, B. Fasel, and J. Luettin, “Fast face detection using MLP and FFT,” in Proceedings of Second International Conference on Audio and Video-based Biometric Person Authentication (AVBPA’99) , 1999, pp. 31–36
1999
Earlier work this paper cites.
K. Minami, H. Nakajima, and T. Toyoshima, “Real-time discrimination of ventricular tachyarrhythmia with Fourier-Transform neural network,” IEEE Transactions on Biomedical Engineering , vol. 46, no. 2, pp. 179–185, 1999
1999
Earlier work this paper cites.
Y.-Q. Zhang and L.-W. Chan, “ForeNet: Fourier recurrent networks for time series prediction,” in Proceedings of International Conference on Neural Information Processing (ICONIP 2000) , 2000, pp. 576–582
2000
Earlier work this paper cites.
H. M. Ozaktas, Z. Zalevsky, and M. A. Kutay, The fractional Fourier transform: with applications in optics and signal processing . Wiley, 2001
2001
Earlier work this paper cites.
H. M. Ozaktas and M. A. Kutay, “The Fractional Fourier Transform,” in 2001 European Control Conference (ECC) . IEEE, 2001, pp. 1477–1483
2001
Earlier work this paper cites.
H. M. El-Bakry and Q. Zhao, “Fast object/face detection using neural networks and Fast Fourier Transform,” International Journal of Signal Processing , vol. 1, no. 3, pp. 182–187, 2004
2004
Earlier work this paper cites.
H. S. Tan, “Fourier neural networks and generalized single hidden layer networks in aircraft engine fault diagnostics,” Journal of Engineering for Gas Turbines and Power , vol. 128, no. 4, pp. 773–782, 2005
2005
Earlier work this paper cites.
W. Zuo, Y. Zhu, and L. Cai, “Fourier-neural-network-based learning control for a class of nonlinear systems with flexible components,” IEEE Transactions on Neural Networks , vol. 20, no. 1, pp. 139–151, 2008
2008
Earlier work this paper cites.
H. Gothwal, S. Kedawat, and R. Kumar, “Cardiac arrhythmias detection in an ECG beat signal using Fast Fourier Transform and artificial neural network,” Journal of Biomedical Science and Engineering , vol. 4, no. 4, pp. 289–296, 2011
2011
Earlier work this paper cites.
Z. Zhang, Y. Wang, and K. Wang, “Fault diagnosis and prognosis using wavelet packet decomposition, Fourier Transform and artificial neural network,” Journal of Intelligent Manufacturing , vol. 24, no. 6, pp. 1213–1227, 2013
2013
Earlier work this paper cites.
S. Liu, “Fourier neural network for machine learning,” in 2013 International Conference on Machine Learning and Cybernetics , vol. 1, 2013, pp. 285–290
2013
Earlier work this paper cites.
M. Mathieu, M. Henaff, and Y. LeCun, “Fast training of convolutional networks through FFTs,” in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings , Y. Bengio and Y. LeCun, Eds., 2014
2014
Earlier work this paper cites.
M. Vetterli, J. Kovačević, and V. K. Goyal, Foundations of Signal Processing . Cambridge University Press, 2014
2014
Earlier work this paper cites.
M. S. Gashler and S. C. Ashmore, “Training deep Fourier neural networks to fit time-series data,” in Intelligent Computing in Bioinformatics . Cham: Springer International Publishing, 2014, pp. 48–55
2014
Earlier work this paper cites.
M. Mathieu, M. Henaff, and Y. LeCun, “Fast training of convolutional networks through FFTs,” in 2nd International Conference on Learning Representations, ICLR , Banff, AB, Canada, 2014
2014
Earlier work this paper cites.
N. Vasilache, J. Johnson, M. Mathieu, S. Chintala, S. Piantino, and Y. LeCun, “Fast convolutional nets with fbfft: A GPU performance evaluation,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
O. Rippel, J. Snoek, and R. P. Adams, “Spectral representations for convolutional neural networks,” in Advances in Neural Information Processing Systems , vol. 28. Curran Associates, Inc., 2015
2015
Earlier work this paper cites.
T. Highlander and A. Rodriguez, “Very efficient training of convolutional neural networks using Fast Fourier Transform and overlap-and-add,” in Proceedings of the British Machine Vision Conference (BMVC) . BMVA Press, 2015, pp. 160.1–160.9
2015
Earlier work this paper cites.
M. Mironovova and J. Bíla, “Fast Fourier Transform for feature extraction and neural network for classification of electrocardiogram signals,” in 2015 Fourth International Conference on Future Generation Communication Technology (FGCT) , 2015, pp. 1–6
2015
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
H. Pratt, B. Williams, F. Coenen, and Y. Zheng, “FCNN: Fourier convolutional neural networks,” in Machine Learning and Knowledge Discovery in Databases . Cham: Springer International Publishing, 2017, pp. 786–798
2017
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP . Brussels, Belgium: Association for Computational Linguistics, 2018, pp. 353–355
2018
Earlier work this paper cites.
J. Zhang, Y. Lin, Z. Song, and I. Dhillon, “Learning long term dependencies via Fourier recurrent units,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80. PMLR, 2018, pp. 5815–5823
2018
Earlier work this paper cites.
J. Ryu, M.-H. Yang, and J. Lim, “DFT-based transformation invariant pooling layer for visual classification,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Radford et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Zhang, Y. Zhao, H. Li, and C. Zong, “Attention with sparsity regularization for neural machine translation and summarization,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 3, pp. 507–518, 2019
2019
Cited alongside, same era.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “MLP-Mixer: An all-MLP architecture for vision,” in Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
T. Nguyen, V. Suliafu, S. Osher, L. Chen, and B. Wang, “FMMformer: Efficient and flexible transformer via decomposed near-field and far-field attention,” in Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
Y. Rao, W. Zhao, Z. Zhu, J. Lu, and J. Zhou, “Global filter networks for image classification,” in Advances in Neural Information Processing Systems , vol. 34. Curran Associates, Inc., 2021, pp. 980–993
2021
Later among the works it cites.
Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. liu, K. Bhattacharya, A. Stuart, and A. Anandkumar, “Fourier neural operator for parametric partial differential equations,” in International Conference on Learning Representations , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Ma, L. Yu, S. Tian, E. Chen, and W. W. Y. Ng, “Global-local mutual attention model for text classification,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 12, pp. 2127–2139, 2019
2019
Cited alongside, same era.
L. Liu, J. Wu, D. Li, L. Senhadji, and H. Shu, “Fractional wavelet scattering network and applications,” IEEE Transactions on Biomedical Engineering , vol. 66, no. 2, pp. 553–563, 2019
2019
Cited alongside, same era.
I. Tenney, D. Das, and E. Pavlick, “BERT rediscovers the classical NLP pipeline,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, 2019, pp. 4593–4601
2019
Cited alongside, same era.
S. H. Khan, M. Hayat, and F. Porikli, “Regularization of deep neural networks with spectral dropout,” Neural Networks , vol. 110, pp. 82–90, 2019
2019
Cited alongside, same era.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A lite BERT for self-supervised learning of language representations,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
W. You, S. Sun, and M. Iyyer, “Hard-coded Gaussian attention for neural machine translation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, 2020, pp. 7689–7700
2020
Cited alongside, same era.
Z. Zheng, S. Huang, R. Weng, X.-Y. Dai, and J. Chen, “Improving self-attention networks with sequential relations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1707–1716, 2020
2020
Cited alongside, same era.
P. Henderson, J. Hu, J. Romoff, E. Brunskill, D. Jurafsky, and J. Pineau, “Towards the systematic reporting of the energy and carbon footprints of machine learning,” Journal of Machine Learning Research , vol. 21, no. 248, pp. 1–43, 2020. [Online]. Available: http://jmlr.org/papers/v21/20-312.html
2020
Cited alongside, same era.
2021
Later among the works it cites.
S. Cao, “Choose a transformer: Fourier or Galerkin,” in Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
J. Gao, T. He, X. Zhou, and S. Ge, “Skeleton-based action recognition with focusing-diffusion graph convolutional networks,” IEEE Signal Processing Letters , vol. 28, pp. 2058–2062, 2021
2021
Later among the works it cites.
N. Chen, S. Watanabe, J. Villalba, P. Żelasko, and N. Dehak, “Non-autoregressive transformer for speech recognition,” IEEE Signal Processing Letters , vol. 28, pp. 121–125, 2021
2021
Later among the works it cites.
H. Xie, Z. Qin, G. Y. Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Transactions on Signal Processing , vol. 69, pp. 2663–2675, 2021
2021
Later among the works it cites.
K. Pratik, B. D. Rao, and M. Welling, “RE-MIMO: Recurrent and permutation equivariant neural MIMO detection,” IEEE Transactions on Signal Processing , vol. 69, pp. 459–473, 2021
2021
Later among the works it cites.
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh, “Nyströmformer: A Nyström-based algorithm for approximating self-attention,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 16, pp. 14 138–14 148, 2021
2021
Later among the works it cites.
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler, “Long range arena : A benchmark for efficient transformers,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=qVyeW-grC2k
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Kaselimi, A. Voulodimos, I. Daskalopoulos, N. Doulamis, and A. Doulamis, “A vision transformer model for convolution-free multilabel classification of satellite imagery in deforestation monitoring,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–9, 2022
2022
Closest in time.
J. An, S. Cho, J. Bang, and M. Kim, “Domain-slot relationship modeling using a pre-trained language encoder for multi-domain dialogue state tracking,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2091–2102, 2022
2022
Closest in time.
C. Wang, S. Dai, Y. Wang, F. Yang, M. Qiu, K. Chen, W. Zhou, and J. Huang, “ARoBERT: An ASR robust pre-trained language model for spoken language understanding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1207–1218, 2022
2022
Closest in time.
2022
Closest in time.
D. Patterson, J. Gonzalez, U. Hölzle, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. R. So, M. Texier, and J. Dean, “The carbon footprint of machine learning training will plateau, then shrink,” Computer , vol. 55, no. 7, pp. 18–28, 2022
2022
Closest in time.
J. Lee-Thorp, J. Ainslie, I. Eckstein, and S. Ontanon, “FNet: Mixing tokens with Fourier Transforms,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Seattle, United States: Association for Computational Linguistics, Jul. 2022, pp. 4296–4313
2022
Closest in time.
X. Zhao, M. Zhang, R. Tao, W. Li, W. Liao, L. Tian, and W. Philips, “Fractional Fourier image transformer for multimodal remote sensing data classification,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–13, 2022
2022
Closest in time.
B. Belainine, F. Sadat, and M. Boukadoum, “End-to-end dialogue generation using a single encoder and a decoder cascade with a multidimension attention mechanism,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–11, 2022
2022
Closest in time.
R. Agrawal, D. Wolff, and S. Dixon, “A convolutional-attentional neural framework for structure-aware performance-score synchronization,” IEEE Signal Processing Letters , vol. 29, pp. 344–348, 2022
2022
Closest in time.
Y. Lu, J. Zhang, J. Zeng, S. Wu, and C. Zong, “Attention analysis and calibration for transformer in natural language generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1927–1938, 2022
2022
Closest in time.
H. Pan, D. Badawi, and A. E. Cetin, “Block Walsh-Hadamard Transform based binary layers in deep neural networks,” ACM Transactions on Embedded Computing Systems , 2022
2022
Closest in time.
C. Nalmpantis, N. Virtsionis Gkalinikis, and D. Vrakas, “Neural Fourier energy disaggregation,” Sensors , vol. 22, no. 2, 2022
2022
Closest in time.
J. Guibas, M. Mardani, Z. Li, A. Tao, A. Anandkumar, and B. Catanzaro, “Efficient token mixing for transformers via adaptive Fourier neural operators,” in International Conference on Learning Representations , 2022
2022
Closest in time.
Z. Liang, Y. Wang, L. Wang, J. Yang, and S. Zhou, “Light field image super-resolution with transformers,” IEEE Signal Processing Letters , vol. 29, pp. 563–567, 2022
2022
Closest in time.
J. Kong, Y. Bian, and M. Jiang, “MTT: Multi-scale temporal transformer for skeleton-based action recognition,” IEEE Signal Processing Letters , vol. 29, pp. 528–532, 2022
2022
Closest in time.
Z. Tian, J. Yi, J. Tao, S. Zhang, and Z. Wen, “Hybrid autoregressive and non-autoregressive transformer models for speech recognition,” IEEE Signal Processing Letters , vol. 29, pp. 762–766, 2022
2022
Closest in time.