Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) with long context capabilities are integral to complex tasks in natural language processing and computational biology, such as text generation and protein sequence analysis.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
Ainslie, J., Ontanon, S., Alberti, C., Cvicek, V., Fisher, Z., Pham, P., Ravula, A., Sanghai, S., Wang, Q., and Yang, L · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Earlier work this paper cites.
Limitations of transformers on clinical text classification
Gao, S., Alawad, M., Young, M. T., Gounley, J., Schaefferkoetter, N., Yoon, H. J., Wu, X.-C., Durbin, E. B., Doherty, J., Stroup, A., et al · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Soft: Softmax-free transformer with linear complexity
Lu, J., Yao, J., Zhang, J., Zhu, X., Xu, H., Gao, W., Xu, C., Xiang, T., and Zhang, L · 2021
Earlier work this paper cites.
Self-attention does not need o ( n 2 ) o(n^{2}) memory
Rabe, M. N. and Staats, C · 2021
Cited alongside, same era.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Clinical-longformer and clinical-bigbird: Transformers for long clinical sequences
Li, Y., Wehbe, R. M., Ahmad, F. S., Wang, H., and Luo, Y · 2022
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al · 2023
Later among the works it cites.
Fingpt: Open-source financial large language models
Yang, H., Liu, X.-Y., and Wang, C. D · 2023
Later among the works it cites.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generative ai and firm values
Eisfeldt, A. L., Schubert, G., and Zhang, M. B · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, L., Rajbhandari, S., and He, Y · 2023
Cited alongside, same era.
From transcripts to insights: Uncovering corporate risks using generative ai
Kim, A., Muhn, M., and Nikolaev, V · 2023
Cited alongside, same era.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2023
Cited alongside, same era.
Large language models in finance: A survey
Li, Y., Wang, S., Ding, H., and Chen, H · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context
Liu, H., Zaharia, M., and Abbeel, P · 2023
Cited alongside, same era.
Climax: A foundation model for weather and climate
Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. K., and Grover, A · 2023
Cited alongside, same era.
Genslms: Genome-scale language models reveal sars-cov-2 evolutionary dynamics
Zvyagin, M., Brace, A., Hippe, K., Deng, Y., Zhang, B., Bohorquez, C. O., Clyde, A., Kale, B., Perez-Rivera, D., Ma, H., et al · 2023
Later among the works it cites.
Bloated disclosures: can chatgpt help investors process information?
Kim, A., Muhn, M., and Nikolaev, V. V · 2024
Closest in time.
Blockwise parallel transformers for large context models
Liu, H. and Abbeel, P · 2024
Closest in time.
Mini-sequence transformer: Optimizing intermediate memory for long sequences training
Luo, C., Zhao, J., Chen, Z., Chen, B., and Anandkumar, A · 2024
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention
Munkhdalai, T., Faruqui, M., and Gopal, S · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
Efficiently training 7b llm with 1 million sequence length on 8 gpus
Zhao, P., Zhang, H., Fu, F., Nie, X., Liu, Q., Yang, F., Peng, Y., Jiao, D., Li, S., Xue, J., et al · 2024
Closest in time.