Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are computationally intensive.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in Neural Information Processing Systems 33 (2020) 1877–1901
1901
Earlier work this paper cites.
2014
Earlier work this paper cites.
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, P. Takala, Good debt or bad debt: Detecting semantic orientations in economic texts, Journal of the Association for Information Science and Technology 65 (4) (2014) 782–796
2014
Earlier work this paper cites.
A. Novikov, D. Podoprikhin, A. Osokin, D. P. Vetrov, Tensorizing neural networks, Advances in Neural Information Processing Systems 28 (2015)
2015
Earlier work this paper cites.
Y. Cheng, F. X. Yu, R. S. Feris, S. Kumar, A. Choudhary, S.-F. Chang, An exploration of parameter redundancy in deep networks with circulant projections, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2857–2865
2015
Earlier work this paper cites.
J. C. S. Alvarado, K. Verspoor, T. Baldwin, Domain adaption of named entity recognition to support credit risk assessment, in: Proceedings of the Australasian Language Technology Association Workshop 2015, 2015, pp. 84–90
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems (2017) 30 (2017)
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al., Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
M. Maia, S. Handschuh, A. Freitas, B. Davis, A. Balahur, Www’18 open challenge: Financial opinion mining and question answering, in: Companion of the The Web Conference 2018, 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9
2019
Earlier work this paper cites.
K. Hayashi, T. Yamaguchi, Y. Sugawara, S.-i. Maeda, Exploring unexplored tensor network decompositions for convolutional neural networks, Advances in Neural Information Processing Systems 32 (2019)
2019
Earlier work this paper cites.
T. Zhang, X.-Y. Liu, cutensor-tubal: Optimized gpu library for low-tubal-rank tensors, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 8583–8587
2019
Earlier work this paper cites.
H. Lu, T. Zhang, X.-Y. Liu, High-performance homomorphic matrix completion on gpus, in: 2019 IEEE 21st International Conference on High Performance Computing and Communications; IEEE 17th International Conference on Smart City; IEEE 5th International Conference on Data Science and Systems (HPCC/SmartCity/DSS), IEEE, 2019, pp. 1627–1634
2019
Earlier work this paper cites.
H. Li, T. Zhang, R. Zhang, X.-Y. Liu, High-performance tensor decoder on gpus for wireless camera networks in iot, in: 2019 IEEE 21st International Conference on High Performance Computing and Communications; IEEE 17th International Conference on Smart City; IEEE 5th International Conference on Data Science and Systems (HPCC/SmartCity/DSS), IEEE, 2019, pp. 1619–1626
2019
Earlier work this paper cites.
T. Zhang, X.-Y. Liu, X. Wang, A. Walid, cutensor-tubal: Efficient primitives for tubal-rank tensor learning operations on gpus, IEEE Transactions on Parallel and Distributed Systems 31 (3) (2019) 595–610
2019
Earlier work this paper cites.
A. Gokaslan, V. Cohen, Openwebtext, https://github.com/Skylion007/openwebtext (2019)
2019
Cited alongside, same era.
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, K. Toutanova, Boolq: Exploring the surprising difficulty of natural yes/no questions, in: NAACL, 2019
2019
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale, 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
T. Zhang, X.-Y. Liu, X. Wang, High performance gpu tensor completion with tubal-sampling pattern, IEEE Transactions on Parallel and Distributed Systems 31 (7) (2020) 1724–1739
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Magic, Twitter financial news senti-ment, http://precog.iiitd.edu.in/people/anupama , (2022)
2022
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, Y. Choi, Piqa: Reasoning about physical commonsense in natural language, in: Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
T. Zhang, W. Kan, X.-Y. Liu, High performance gpu primitives for graph-tensor learning operations, Journal of Parallel and Distributed Computing 148 (2021) 125–137
2021
Cited alongside, same era.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al., LoRA: Low-rank adaptation of large language models, in: International Conference on Learning Representations, 2022
2022
Cited alongside, same era.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, J. Tang, Glm: General language model pretraining with autoregressive blank infilling, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 320–335
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., Training language models to follow instructions with human feedback, Advances in Neural Information Processing Systems 35 (2022) 27730–27744
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, E. Goffinet, D. Heslow, J. Launay, Q. Malartic, B. Noune, B. Pannier, G. Penedo, Falcon-40B: an open large language model with state-of-the-art performance (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, S. Han, Smoothquant: Accurate and efficient post-training quantization for large language models, in: International Conference on Machine Learning, PMLR, 2023, pp. 38087–38099
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, I. Stoica, Efficient memory management for large language model serving with pagedattention, in: Proceedings of the 29th Symposium on Operating Systems Principles, 2023, pp. 611–626
2023
Later among the works it cites.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, T. B. Hashimoto, Stanford alpaca: An instruction-following llama model (2023)
2023
Later among the works it cites.
X.-Y. Liu, G. Wang, H. Yang, D. Zha, Data-centric FinGPT: Democratizing internet-scale data for financial large language models, NeurIPS Workshop on Instruction Tuning and Instruction Following (2023)
2023
Later among the works it cites.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, Y. Liu, Roformer: Enhanced transformer with rotary position embedding, Neurocomputing 568 (2024) 127063
2024
Closest in time.