Fetching the paper…
Reading the bibliography…
Despite recent advances in subquadratic attention mechanisms or state-space models, processing long token sequences still imposes significant computational requirements.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Earlier work this paper cites.
Graph transformer networks
Yun, S., Jeong, M., Kim, R., Kang, J., and Kim, H. J · 2019
Earlier work this paper cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination
Goyal, S., Choudhury, A. R., Raje, S., Chakaravarthy, V., Sabharwal, Y., and Verma, A · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Monash time series forecasting archive
Godahewa, R. W., Bergmeir, C., Webb, G. I., Hyndman, R., and Montero-Manso, P · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
Token pooling in vision transformers
Marin, D., Chang, J.-H. R., Ranjan, A., Prabhu, A., Rastegari, M., and Tuzel, O · 2021
Earlier work this paper cites.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Wu, H., Xu, J., Wang, J., and Long, M · 2021
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W · 2021
Earlier work this paper cites.
Cirstea, R.-G., Guo, C., Yang, B., Kieu, T., Dong, X., and Pan, S · 2022
Earlier work this paper cites.
Hebo: Pushing the limits of sample-efficient hyperparameter optimisation
Cowen-Rivers, A., Lyu, W., Tutunov, R., Wang, Z., Grosnit, A., Griffiths, R.-R., Maravel, A., Hao, J., Wang, J., Peters, J., and Bou Ammar, H · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Re, C · 2022
Earlier work this paper cites.
Adavit: Adaptive vision transformers for efficient image recognition
Meng, L., Li, H., Chen, B.-C., Lan, S., Wu, Z., Jiang, Y.-G., and Lim, S.-N · 2022
Cited alongside, same era.
Etsformer: Exponential smoothing transformers for time-series forecasting
Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S · 2022
Cited alongside, same era.
Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting
Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R · 2022
Cited alongside, same era.
Thop: Pytorch-opcounter
Zhu, L · 2022
Cited alongside, same era.
Token merging for fast stable diffusion
Bolya, D. and Hoffman, J · 2023
Cited alongside, same era.
Token merging: Your vit but faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Itransformer: Inverted transformers are effective for time series forecasting
Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M · 2023
Later among the works it cites.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Nguyen, E., Poli, M., Faizi, M., Thomas, A., Wornow, M., Birch-Sykes, C., Massaroli, S., Patel, A., Rabideau, C., Bengio, Y., Ermon, S., Ré, C., and Baccus, S · 2023
Later among the works it cites.
A time series is worth 64 words: Long-term forecasting with transformers
Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Re, C · 2023
Later among the works it cites.
Lag-llama: Towards foundation models for probabilistic time series forecasting
Rasul, K., Ashok, A., Williams, A. R., Ghonia, H., Bhagwatkar, R., Khorasani, A., Bayazi, M. J. D., Adamopoulos, G., Riachi, R., Hassen, N., Biloš, M., Garg, S., Schneider, A., Chapados, N., Drouin, A., Zantedeschi, V., Nevmyvaka, Y., and Rish, I · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learned thresholds token merging and pruning for vision transformers
Bonnaerens, M. and Dambre, J · 2023
Cited alongside, same era.
Diffrate: Differentiable compression rate for efficient vision transformers
Chen, M., Shao, W., Xu, P., Lin, M., Zhang, K., Chao, F., Ji, R., Qiao, Y., and Luo, P · 2023
Cited alongside, same era.
A decoder-only foundation model for time-series forecasting
Das, A., Kong, W., Sen, R., and Zhou, Y · 2023
Cited alongside, same era.
Genomic benchmarks: a collection of datasets for genomic sequence classification
Grešová, K., Martinek, V., Čechák, D., Šimeček, P., and Alexiou, P · 2023
Cited alongside, same era.
Large language models are zero-shot time series forecasters
Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G · 2023
Cited alongside, same era.
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Are transformers effective for time series forecasting?
Zeng, A., Chen, M., Zhang, L., and Xu, Q · 2023
Later among the works it cites.
Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting
Zhang, Y. and Yan, J · 2023
Later among the works it cites.
One fits all: Power general time series analysis by pretrained lm
Zhou, T., Niu, P., wang, x., Sun, L., and Jin, R · 2023
Later among the works it cites.
Chronos: Learning the language of time series
Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., Zschiegner, J., Maddix, D. C., Wang, H., Mahoney, M. W., Torkkola, K., Wilson, A. G., Bohlke-Schneider, M., and Wang, Y · 2024
Closest in time.
Token fusion: Bridging the gap between token pruning and token merging
Kim, M., Gao, S., Hsu, Y.-C., Shen, Y., and Jin, H · 2024
Closest in time.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
Raposo, D., Ritter, S., Richards, B., Lillicrap, T., Humphreys, P. C., and Santoro, A · 2024
Closest in time.
Accelerating transformers with spectrum-preserving token merging, 2024
Tran, H.-C., Nguyen, D. M. H., Nguyen, D. M., Nguyen, T.-T., Le, N., Xie, P., Sonntag, D., Zou, J. Y., Nguyen, B. T., and Niepert, M · 2024
Closest in time.
Unified training of universal time series forecasting transformers
Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D · 2024
Closest in time.