Fetching the paper…
Reading the bibliography…
Artificial intelligence (AI) methods have become critical in scientific applications to help accelerate scientific discovery.
2005
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Mathuriya, D. Bard, P. Mendygral, L. Meadows, J. Arnemann, L. Shao, S. He, T. Kärnä, D. Moise, S. J. Pennycook, K. J. Maschhoff, J. Sewall, N. Kumar, S. Ho, M. F. Ringenburg, Prabhat, and V. W. Lee, “CosmoFlow: Using deep learning to learn the universe at scale,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2018, Dallas, TX, USA, November 11-16, 2018 . IEEE / ACM, 2018, pp. 65:1–65:11
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018. [Online]. Available: https://paperswithcode.com/paper/improving-language-understanding-by
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Gokaslan and V. Cohen, “OpenWebText Corpus,” http://Skylion007.github.io/OpenWebTextCorpus , 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko et al. , “Highly accurate protein structure prediction with AlphaFold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Earlier work this paper cites.
K. Tunyasuvunakool, J. Adler, Z. Wu, T. Green, M. Zielinski, A. Žídek, A. Bridgland, A. Cowie, C. Meyer, A. Laydon et al. , “Highly accurate protein structure prediction for the human proteome,” Nature , vol. 596, no. 7873, pp. 590–596, 2021
2021
Earlier work this paper cites.
J. Yin, A. Tsaris, S. Dash, R. Miller, F. Wang, and M. A. Shankar, “Comparative evaluation of deep learning workloads for leadership-class systems,” BenchCouncil Transactions on Benchmarks, Standards and Evaluations , vol. 1, no. 1, p. 100005, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2772485921000053
2021
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on GPU clusters using Megatron-LM,” in Proceedings of SC’21 , 2021
2021
Earlier work this paper cites.
M. Zvyagin, A. Brace, K. Hippe, Y. Deng, B. Zhang, C. O. Bohorquez, A. Clyde, B. Kale, D. Perez-Rivera, H. Ma et al. , “GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics,” bioRxiv , pp. 2022–10, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
“LLM on Sambanova Datascale,” https://sambanova.ai/blog/achieving-best-in-class-large-language-model-accuracy-in-low-resource-settings/ , 2022
2022
Cited alongside, same era.
M. Emani, Z. Xie, S. Raskar, V. Sastry, W. Arnold, B. Wilson, R. Thakur, V. Vishwanath, Z. Liu, M. E. Papka et al. , “A comprehensive evaluation of novel AI accelerators for deep learning workloads,” in Proceedings of PMBS Workshop 2022 , 2022
2022
Cited alongside, same era.
“GenSLMs: Genome-scale language models,” https://github.com/ramanathanlab/genslm , 2022
2022
Cited alongside, same era.
“Habana Gaudi2- Memory Efficient training,” https://developer.habana.ai/blog/memory-efficient-training-on-habana-gaudi-with-deepspeed/ , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” in Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
“Cerebras Weight Streaming,” https://www.cerebras.net/blog/linear-scaling-made-possible-with-weight-streaming , September 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“SambaNova Suite GPT,” https://sambanova.ai/solutions/gpt/ , 2023
2023
Cited alongside, same era.
“Cerebras-GPT,” https://www.cerebras.net/blog/cerebras-gpt-a-family-of-open-compute-efficient-large-language-models// , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“Graphcore GPT-J,” https://www.graphcore.ai/posts/fine-tuned-gpt-j-a-cost-effective-alternative-to-gpt-4-for-nlp-tasks , 2023
2023
Cited alongside, same era.
2023
Closest in time.
J. Yin, S. Dash, J. Gounley, F. Wang, and G. Tourassi, “Evaluation of pre-training large language models on leadership-class supercomputers,” The Journal of Supercomputing , pp. 1–22, 06 2023
2023
Closest in time.
OpenAI, “GPT-4 technical report,” 2023. [Online]. Available: https://arxiv.org/abs/2303.08774
2023
Closest in time.
“Polaris supercomputing system,” https://www.alcf.anl.gov/polaris , August 2023
2023
Closest in time.
“Weight Streaming Mode,” https://docs.cerebras.net/en/latest/wsc/cerebras-basics/cerebras-execution-modes.html , August 2023
2023
Closest in time.
2023
Closest in time.
“Nvidia NeMo Framework,” https://developer.nvidia.com/nemo , 2023
2023
Closest in time.
“Control numerical precision level,” https://docs.cerebras.net/en/latest/wsc/how_to_guides/cs-1-data-formats.html , August 2023
2023
Closest in time.
“IPU Replication factor,” https://docs.graphcore.ai/projects/popart-user-guide/en/latest/glossary.html#term-Replication-factor , August 2023
2023
Closest in time.