Fetching the paper…
Reading the bibliography…
The advent of Transformers marked a significant breakthrough in sequence modelling, providing a highly performant architecture capable of leveraging GPU parallelism.
Parallel prefix computation
Ladner, R. E. and Fischer, M. J. (1980) · 1980
Earlier work this paper cites.
Data parallel algorithms
Hillis, W. D. and Steele, G. L. (1986) · 1986
Earlier work this paper cites.
Prefix sums and their applications
Blelloch, G. E. (1990) · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Cho, K., Van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
The uea multivariate time series classification archive, 2018
Bagnall, A., Dau, H. A., Lines, J., Flynn, M., Large, J., Bostrom, A., Southam, P., and Keogh, E. (2018) · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. (2020) · 2020
Earlier work this paper cites.
Self-attentive hawkes process
Zhang, Q., Lipani, A., Kirnap, O., and Yilmaz, E. (2020) · 2020
Cited alongside, same era.
Zuo, S., Jiang, H., Li, Z., Zhao, T., and Zha, H. (2020) · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J. (2021) · 2021
Cited alongside, same era.
Minimal implementation of decision transformer
Barhate, N. (2022) · 2022
Cited alongside, same era.
What is time series classification?
Dinger, T., Chang, Y.-c., Pavuluri, R., and Subramanian, S. (2022) · 2022
Transformers in reinforcement learning: A survey
Agarwal, P., Rahman, A. A., St-Charles, P.-L., Prince, S. J., and Kahou, S. E. (2023) · 2023
Later among the works it cites.
Meta temporal point processes
Bae, W., Ahmed, M. O., Tung, F., and Oliveira, G. L. (2023) · 2023
Later among the works it cites.
Memory efficient neural processes via constant memory attention block
Feng, L., Tung, F., Hajimirsadeghi, H., Bengio, Y., and Ahmed, M. O. (2023) · 2023
Later among the works it cites.
A comprehensive survey on applications of transformers for deep learning tasks
Islam, S., Elmekki, H., Elsebai, A., Bentahar, J., Drawel, N., Rjoub, G., and Pedrycz, W. (2023) · 2023
Later among the works it cites.
Transformers in speech processing: A survey
Latif, S., Zaidi, A., Cuayahuitl, H., Shamshad, F., Shoukat, M., and Qadir, J. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of transformers
Lin, T., Wang, Y., Liu, X., and Qiu, X. (2022) · 2022
Cited alongside, same era.
Non-stationary transformers: Exploring the stationarity in time series forecasting
Liu, Y., Wu, H., Wang, J., and Long, M. (2022) · 2022
Cited alongside, same era.
Self-attention does not need o ( n 2 ) o(n^{2}) memory
Rabe, M. N. and Staats, C. (2022) · 2022
Cited alongside, same era.
Transformers in time series: A survey
Wen, Q., Zhou, T., Zhang, C., Chen, W., Ma, Z., Yan, J., and Sun, L. (2022) · 2022
Cited alongside, same era.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A. (2022) · 2022
Cited alongside, same era.
Later among the works it cites.
A survey on transformers in reinforcement learning
Li, W., Luo, H., Lin, Z., Zhang, C., Lu, Z., and Ye, D. (2023) · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Derczynski, L., et al. (2023) · 2023
Later among the works it cites.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J. (2023) · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F. (2023) · 2023
Later among the works it cites.
Timesnet: Temporal 2d-variation modeling for general time series analysis
Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. (2023) · 2023
Later among the works it cites.
Empowering time series analysis with large language models: A survey
Jiang, Y., Pan, Z., Zhang, X., Garg, S., Schneider, A., Nevmyvaka, Y., and Song, D. (2024) · 2024
Closest in time.