Fetching the paper…
Reading the bibliography…
Transformer-based models have emerged as one of the most widely used architectures for natural language processing, natural language generation, and image generation.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009
2009
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of naacL-HLT , vol. 1, 2019, p. 2
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
J. Choquette, E. Lee, R. Krashinsky, V. Balan, and B. Khailany, “3.2 the a100 datacenter gpu and ampere architecture,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IEEE, 2021, pp. 48–50
2021
Earlier work this paper cites.
A. Ivanov, N. Dryden, T. Ben-Nun, S. Li, and T. Hoefler, “Data movement is all you need: A case study on optimizing transformers,” Proceedings of Machine Learning and Systems , vol. 3, pp. 711–732, 2021
2021
Earlier work this paper cites.
Z. Jia and P. Van Sandt, “Dissecting the ampere gpu architecture via microbenchmarking,” in GPU Technology Conference , 2021
2021
Earlier work this paper cites.
J. R. Stevens, R. Venkatesan, S. Dai, B. Khailany, and A. Raghunathan, “Softermax: Hardware/software co-design of an efficient softmax for transformers,” in 2021 58th ACM/IEEE Design Automation Conference (DAC) . IEEE, 2021, pp. 469–474
2021
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for Transformer-Based generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . Carlsbad, CA: USENIX Association, Jul. 2022, pp. 521–538. [Online]. Available: https://www.usenix.org/conference/osdi22/presentation/yu
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“GitHub - NVIDIA/cutlass: CUDA Templates for Linear Algebra Subroutines — github.com,” https://github.com/NVIDIA/cutlass , [Accessed 01-04-2024]
2024
Closest in time.
“Introducing ChatGPT — openai.com,” https://openai.com/blog/chatgpt , [Accessed 01-04-2024]
2024
Closest in time.
“Introducing the next generation of claude,” https://www.anthropic.com/news/claude-3-family , [Accessed 19-04-2024]
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Osama, D. Merrill, C. Cecka, M. Garland, and J. D. Owens, “Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,” in Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming , 2023, pp. 429–431
2023
Cited alongside, same era.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“Cute layouts.” https://github.com/NVIDIA/cutlass/blob/main/media/docs/cute/01_layout.md , [Accessed 19-04-2024]
2024
Cited alongside, same era.
“Cute tensors.” https://github.com/NVIDIA/cutlass/blob/main/media/docs/cute/03_tensor.md. , [Accessed 19-04-2024]
2024
Cited alongside, same era.
“Cute’s support for matrix multiply-accumulate instructions.” https://github.com/NVIDIA/cutlass/blob/main/media/docs/cute/0t_mma_atom.md , [Accessed 19-04-2024]
2024
Cited alongside, same era.
2024
Closest in time.
K. Hong, G. Dai, J. Xu, Q. Mao, X. Li, J. Liu, Y. Dong, Y. Wang et al. , “Flashdecoding++: Faster large language model inference with asynchronization, flat gemm optimization, and heuristics,” Proceedings of Machine Learning and Systems , vol. 6, pp. 148–161, 2024
2024
Closest in time.
G. Inc., “An important next step on our AI journey — blog.google,” https://blog.google/technology/ai/bard-google-ai-search-updates/ , [Accessed 31-03-2024]
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Spataro and M. Inc., “Introducing Microsoft 365 Copilot – your copilot for work - The Official Microsoft Blog — blogs.microsoft.com,” https://blogs.microsoft.com/blog/2023/03/16/introducing-microsoft-365-copilot/-your-copilot-for-work/ , [Accessed 31-03-2024]
2024
Closest in time.
Z. Ye, L. Chen, R. Lai, Y. Zhao, S. Zheng, J. Shao, B. Hou, H. Jin, Y. Zuo, L. Yin, T. Chen, and L. Ceze, “Accelerating self-attentions for llm serving with flashinfer,” February 2024. [Online]. Available: https://flashinfer.ai/2024/02/02/introduce-flashinfer.html
2024
Closest in time.