Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications.
X. Li and D. Roth, “Learning question classifiers,” in COLING 2002: The 19th International Conference on Computational Linguistics , 2002
2002
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, “Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2017, pp. 1601–1611
2017
Earlier work this paper cites.
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning, “Hotpotqa: A dataset for diverse, explainable multi-hop question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018, pp. 2369–2380
2018
Earlier work this paper cites.
W. He, K. Liu, J. Liu, Y. Lyu, S. Zhao, X. Xiao, Y. Liu, Y. Wang, H. Wu, Q. She et al. , “Dureader: a chinese machine reading comprehension dataset from real-world applications,” ACL 2018 , p. 37, 2018
2018
Earlier work this paper cites.
T. Kočiskỳ, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette, “The narrativeqa reading comprehension challenge,” Transactions of the Association for Computational Linguistics , vol. 6, pp. 317–328, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
B. Gliwa, I. Mochol, M. Biesek, and A. Wawer, “Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,” EMNLP-IJCNLP 2019 , p. 70, 2019
2019
Earlier work this paper cites.
A. R. Fabbri, I. Li, T. She, S. Li, and D. Radev, “Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 1074–1084
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
X. Ho, A.-K. D. Nguyen, S. Sugawara, and A. Aizawa, “Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps,” in Proceedings of the 28th International Conference on Computational Linguistics , 2020, pp. 6609–6625
2020
Earlier work this paper cites.
J. Yin, A. Tsaris, S. Dash, R. Miller, F. Wang, and M. A. Shankar, “Comparative evaluation of deep learning workloads for leadership-class systems,” BenchCouncil Transactions on Benchmarks, Standards and Evaluations , vol. 1, no. 1, p. 100005, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2772485921000053
2021
Earlier work this paper cites.
M. Emani, V. Vishwanath, C. Adams, M. E. Papka, R. Stevens, L. Florescu, S. Jairath, W. Liu, T. Nama, and A. Sujeeth, “Accelerating scientific applications with sambanova reconfigurable dataflow architecture,” Computing in Science & Engineering , vol. 23, no. 2, pp. 114–119, 2021
2021
Earlier work this paper cites.
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner, “A dataset of information-seeking questions and answers anchored in research papers,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 4599–4610
2021
Earlier work this paper cites.
L. Huang, S. Cao, N. Parulian, H. Ji, and L. Wang, “Efficient attentions for long document summarization,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 1419–1436
2021
Earlier work this paper cites.
M. Zhong, D. Yin, T. Yu, A. Zaidi, M. Mutuma, R. Jha, A. Hassan, A. Celikyilmaz, Y. Liu, X. Qiu et al. , “Qmsum: A new benchmark for query-based multi-domain meeting summarization,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 5905–5921
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Emani, Z. Xie, S. Raskar, V. Sastry, W. Arnold, B. Wilson, R. Thakur, V. Vishwanath, Z. Liu, M. E. Papka, C. O. Bohorquez, R. Weisner, K. Li, Y. Sheng, Y. Du, J. Zhang, A. Tsyplikhin, G. Khaira, J. Fowers, R. Sivakumar, V. Godsoe, A. Macias, C. Tekur, and M. Boyd, “A comprehensive evaluation of novel ai accelerators for deep learning workloads,” in 2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS) , 2022, pp. 13–25
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Earlier work this paper cites.
Nvidia, “Pynvml,” 2022. [Online]. Available: https://pypi.org/project/pynvml
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Earlier work this paper cites.
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A survey of quantization methods for efficient neural network inference,” in Low-Power Computer Vision . Chapman and Hall/CRC, 2022, pp. 291–326
2022
Earlier work this paper cites.
A. Kuzmin, M. Van Baalen, Y. Ren, M. Nagel, J. Peters, and T. Blankevoort, “Fp8 quantization: The power of the exponent,” Advances in Neural Information Processing Systems , vol. 35, pp. 14 651–14 662, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
K. T. Chitty-Venkata and A. K. Somani, “Neural architecture search survey: A hardware perspective,” ACM Computing Surveys , vol. 55, no. 4, pp. 1–36, 2022
2022
Earlier work this paper cites.
K. T. Chitty-Venkata, M. Emani, V. Vishwanath, and A. K. Somani, “Neural architecture search for transformers: A survey,” IEEE Access , vol. 10, pp. 108 374–108 412, 2022
2022
Earlier work this paper cites.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in International conference on machine learning . PMLR, 2022, pp. 18 332–18 346
2022
Cited alongside, same era.
2022
Cited alongside, same era.
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “Musique: Multihop questions via single-hop question composition,” Transactions of the Association for Computational Linguistics , vol. 10, pp. 539–554, 2022
2022
Cited alongside, same era.
K. T. Chitty-Venkata, S. Mittal, M. Emani, V. Vishwanath, and A. K. Somani, “A survey of techniques for optimizing transformer inference,” Journal of Systems Architecture , p. 102990, 2023
2023
2024
Closest in time.
2024
Closest in time.
Y. Park, K. Budhathoki, L. Chen, J. M. Kübler, J. Huang, M. Kleindessner, J. Huan, V. Cevher, Y. Wang, and G. Karypis, “Inference optimization of foundation models on ai accelerators,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 6605–6615
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
mistralai, “Mistral-7b-v0.1,” 2023. [Online]. Available: https://huggingface.co/mistralai/Mistral-7B-v0.1
2023
Cited alongside, same era.
J. Yin, S. Dash, J. Gounley, F. Wang, and G. Tourassi, “Evaluation of pre-training large language models on leadership-class supercomputers,” The Journal of Supercomputing , pp. 1–22, 06 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Choquette, “Nvidia hopper h100 gpu: Scaling performance,” IEEE Micro , vol. 43, no. 3, pp. 9–17, 2023
2023
Cited alongside, same era.
Nvidia, “Tensor-rt llm,” 2023. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
2023
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Cited alongside, same era.
Microsoft, “Deepspeed mii,” 2023. [Online]. Available: https://github.com/microsoft/DeepSpeed-MII
2023
Cited alongside, same era.
M. Emani, S. Foreman, V. Sastry, Z. Xie, S. Raskar, W. Arnold, R. Thakur, V. Vishwanath, M. E. Papka, S. Shanmugavelu, D. Gandhi, H. Zhao, D. Ma, K. Ranganath, R. Weisner, J. Chen, Y. Yang, N. Vassilieva, B. C. Zhang, S. Howland, and A. Tsyplikhin, “Toward a holistic performance evaluation of large language models across diverse ai accelerators,” in 2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) . Los Alamitos, CA, USA: IEEE Computer Society, may 2024, pp. 1–10. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/IPDPSW63119.2024.00016
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Corporation, “Nvidia a100 tensor core gpu architecture,” Nvidia, White Paper, 2023, accessed: 2024-07-21. [Online]. Available: https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia-ampere-architecture-whitepaper.pdf
2024
Closest in time.
N. Corporation, “Nvidia gh200 grace hopper superchip architecture,” Nvidia, White Paper, 2023, accessed: 2024-07-21. [Online]. Available: https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-hopper?ncid=no-ncid
2024
Closest in time.
I. Advanced Micro Devices, “Amd cdna 2 architecture,” AMD, White Paper, 2023, accessed: 2024-07-21. [Online]. Available: https://www.amd.com/content/dam/amd/en/documents/instinct-business-docs/white-papers/amd-cdna2-white-paper.pdf
2024
Closest in time.
——, “Amd cdna 3 architecture,” AMD, White Paper, 2024
2024
Closest in time.
I. Corporation, “Habana gaudi 2 white paper,” Intel, White Paper, 2023, accessed: 2024-07-21. [Online]. Available: https://www.intel.com/content/www/us/en/content-details/784827/gaudi-2-white-paper.html
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Corporation, “Nvidia h100 tensor core gpu architecture,” Nvidia, White Paper, 2023, accessed: 2024-07-21. [Online]. Available: https://resources.nvidia.com/en-us-tensor-core?ncid=no-ncid
2024
Closest in time.
2024
Closest in time.
Qwen, “Qwen2-1.5b,” 2024. [Online]. Available: https://huggingface.co/Qwen/Qwen2-1.5B
2024
Closest in time.
SambaNova, “Sambastudio,” 2024. [Online]. Available: https://docs.sambanova.ai/sambastudio/latest/sambastudio-intro.html
2024
Closest in time.
meta llama, “Meta-llama-3-8b,” 2024. [Online]. Available: https://huggingface.co/meta-llama/Meta-Llama-3-8B
2024
Closest in time.
Qwen, “Qwen2-57b-a14b,” 2024. [Online]. Available: https://huggingface.co/Qwen/Qwen2-57B-A14B
2024
Closest in time.
S. A. Research, “Snowflake arctic: The best llm for enterprise ai — efficiently intelligent, truly open,” 2024. [Online]. Available: https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/
2024
Closest in time.
2024
Closest in time.
Qwen, “Qwen2-7b,” 2024. [Online]. Available: https://huggingface.co/Qwen/Qwen2-7B
2024
Closest in time.
——, “Meta-llama-3-70b,” 2024. [Online]. Available: https://huggingface.co/meta-llama/Meta-Llama-3-70B
2024
Closest in time.
Qwen, “Qwen2-72b,” 2024. [Online]. Available: https://huggingface.co/Qwen/Qwen2-72B
2024
Closest in time.
Meta, “Building meta’s genai infrastructure,” 2024. [Online]. Available: https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/
2024
Closest in time.