Fetching the paper…
Reading the bibliography…
Long-range sequence modeling is a crucial aspect of natural language processing and time series analysis.
J. L. Elman, “Finding structure in time,” Cognitive Science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, “Recurrent neural network based language model,” Interspeech , vol. 2, pp. 1045–1048, 2010
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” arXiv preprint arXiv:1410.3916 , 2014
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwińska, S. G. Colmenarejo, J. Światkowski, D. Tan, S. Mohamed et al. , “Hybrid computing using a neural network with dynamic external memory,” Nature , vol. 538, no. 7626, pp. 471–476, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” OpenAI Blog , 2018
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
M. Tan and S. Liu, “Hybrid memory networks for sequential data processing,” ICML , 2022
2022
Later among the works it cites.
Y. Tay, M. Dehghani et al. , “Synthformer: Beyond token-level self-attention,” NeurIPS , 2022
2022
Later among the works it cites.
A. Gu and T. Dao, “Scaling structured state space models for long sequences,” NeurIPS , 2022
2022
Later among the works it cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling models with mixture-of-experts layers,” Nature Machine Intelligence , 2022
2022
Later among the works it cites.
R. Dutta et al. , “Neural memory architectures for sequential decision-making,” ICLR , 2023
2023
Later among the works it cites.
H. Liu and H. Rao, “Recurrent attention mechanisms for efficient long-sequence processing,” ICML , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “Hippo: Recurrent memory with optimal polynomial projections,” Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Tran and K. Zhou, “Rnn++: A lightweight, optimized recurrent neural network for long sequences,” NeurIPS , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Z. Chen and T. Wang, “Dynamic memory transformers for efficient long-sequence modeling,” ACL , 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Kim and J. Park, “Modular mixture of experts for efficient sequence modeling,” ICML , 2023
2023
Later among the works it cites.
2024
Later among the works it cites.
Y. Wang et al. , “Compact state space models for efficient long-term dependencies,” ICLR , 2024
2024
Later among the works it cites.