Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown remarkable capabilities in various tasks.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , Language models are few-shot learners, Advances in neural information processing systems , 2020, 33
1901
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser and I. Polosukhin, Attention is all you need, Advances in neural information processing systems , 2017, 30
2017
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam and D. Kalenichenko, Quantization and training of neural networks for efficient integer-arithmetic-only inference , Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2704–2713
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , Language models are unsupervised multitask learners, OpenAI blog , 2019, 1
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li and P. J. Liu, Exploring the limits of transfer learning with a unified text-to-text transformer, The Journal of Machine Learning Research , 2020, 21
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li and Y. He, Zeroquant: Efficient and affordable post-training quantization for large-scale transformers, Advances in Neural Information Processing Systems , 2022, 35
2022
Cited alongside, same era.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth and S. Han, Smoothquant: Accurate and efficient post-training quantization for large language models , International Conference on Machine Learning, 2023, pp. 38087–38099
2023
Cited alongside, same era.
G. Orlanski, K. Xiao, X. Garcia, J. Hui, J. Howland, J. Malmaud, J. Austin, R. Singh and M. Catasta, Measuring the impact of programming language distribution , International Conference on Machine Learning, 2023, pp. 26619–26645
2023
Closest in time.
2023
Closest in time.
R. OpenAI, GPT-4 technical report. arXiv 2303.08774, View in Article , 2023
2023
Closest in time.