Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable abilities in natural language processing.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer Sentinel Mixture Models,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root Mean Square Layer Normalization,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019
2019
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The Curious Case of Neural Text Degeneration,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
N. Shazeer, “GLU Variants Improve Transformer,” arXiv preprint arXiv:2002.05202, 2020
2020
Earlier work this paper cites.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2020
2020
Earlier work this paper cites.
T. B. Brown et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, G. Li, and J. Cheng, “Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language Processing,” in Proceedings of the 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE) , 2021
2021
Cited alongside, same era.
2022
Cited alongside, same era.
C. Peng, X. Yang, A. Chen, et al., “A Study of Generative Large Language Model for Medical Research and Healthcare,” npj Digital Medicine, vol. 6, no. 210, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Pan, H. Cao, Y. Zhu, J. Liu, and B. Li, “Contextual Client Selection for Efficient Federated Learning Over Edge Devices,” IEEE Transactions on Mobile Computing , vol. 23, no. 6, pp. 6538–6548, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
X. Liu et al. , “Influence Pathway Discovery on Social Media,” in 2023 IEEE 9th International Conference on Collaboration and Internet Computing (CIC) , 2023
2023
Cited alongside, same era.
S. Li, R. Zhao, M. Li, H. Ji, C. Callison-Burch, and J. Han, “Open-Domain Hierarchical Event Schema Induction by Incremental Prompting and Verification,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023
2023
Cited alongside, same era.
Y. Liu, J. Huang, and K. Chang, “Ask To The Point: Open-Domain Entity-Centric Question Generation,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023
2023
Cited alongside, same era.
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Chen, J. Zhang, Y. Du, S. Xiang, Z. Yue, N. Zhang, Y. Cai, and Z. Zhang, “Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference,” ACM Transactions on Reconfigurable Technology and Systems , 2024
2024
Closest in time.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI blog. [Online]. Available: https://openai.com/research/better-language-models . Accessed: Mar. 01, 2024
2024
Closest in time.
S. Zeng et al., “FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs,” in Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays , 2024
2024
Closest in time.