Fetching the paper…
Reading the bibliography…
In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing (NLP) tasks, such as question-answering, sentiment analysis, text summarization, and machine translation.
B. Goertzel and C. Pennachin, Artificial general intelligence . Springer, 2007, vol. 2
2007
Earlier work this paper cites.
T. G. Kolda and B. W. Bader, “Tensor decompositions and applications,” SIAM review , vol. 51, no. 3, pp. 455–500, 2009
2009
Earlier work this paper cites.
I. V. Oseledets, “Tensor-train decomposition,” SIAM Journal on Scientific Computing , vol. 33, no. 5, pp. 2295–2317, 2011
2011
Earlier work this paper cites.
A. Novikov et al. , “Tensorizing neural networks,” in Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
A. Novikov, D. Podoprikhin, A. Osokin, and D. P. Vetrov, “Tensorizing neural networks,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
L. Li, K. G. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: Bandit-based configuration evaluation for hyperparameter optimization.” in ICLR (Poster) , 2017, p. 53
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
X. Ma et al. , “A tensorized transformer for language modeling,” in Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International conference on machine learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Superglue: A stickier benchmark for general-purpose language understanding systems,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2023
Later among the works it cites.
OpenAI et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Bhattarai et al. , “Distributed non-negative tensor train decomposition,” in 2020 IEEE HPEC . IEEE, 2020, pp. 1–7
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
M. U. Hadi, K. Y. Azeez, U. F. Mohammed, M. H. Farag, H. A. Omar, and M. Ahmed, “Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects,” Authorea Preprints , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Ding, Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.-M. Chan, W. Chen et al. , “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nature Machine Intelligence , vol. 5, no. 3, pp. 220–235, 2023
2023
Later among the works it cites.
Meta, “Introducing llama: A foundational, 65-billion-parameter large language model,” 2023, https://ai.meta.com/blog/large-language-model-llama-meta-ai/
2023
Later among the works it cites.
Ray, “How to use tune with pytorch,” 2023, https://docs.ray.io/en/latest/tune/examples/tune-pytorch-cifar.html
2023
Later among the works it cites.