Fetching the paper…
Reading the bibliography…
This paper presents Fauno, the first and largest open-source Italian conversational Large Language Model (LLM).
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., · 1901
Earlier work this paper cites.
Bart-it: An efficient sequence-to-sequence model for italian text summarization,
M. La Quatra, L. Cagliero, · 1999
Earlier work this paper cites.
The wacky wide web: a collection of very large linguistically processed web-crawled corpora,
M. Baroni, S. Bernardini, A. Ferraresi, E. Zanchetta, · 2009
Earlier work this paper cites.
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, · 2018
Earlier work this paper cites.
Self-attentive sequential recommendation,
W.-C. Kang, J. McAuley, · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., · 2019
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation,
Y. Zhang, S. Sun, M. Galley, Y.-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, B. Dolan, · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP,
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, S. Gelly, · 2019
Earlier work this paper cites.
A question-entailment approach to question answering,
A. Ben Abacha, D. Demner-Fushman, · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., · 2020
Earlier work this paper cites.
Jukebox: A generative model for music,
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, I. Sutskever, · 2020
Earlier work this paper cites.
Towards precise completion of deformable shapes,
O. Halimi, I. Imanuel, O. Litany, G. Trappolini, E. Rodolà, L. Guibas, R. Kimmel, · 2020
Earlier work this paper cites.
Towards a human-like open-domain chatbot,
D. Adiwardana, M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha, G. Nemade, Y. Lu, et al., · 2020
Cited alongside, same era.
Geppetto carves italian into a language model,
L. D. Mattei, M. Cafagna, F. Dell’Orletta, M. Nissim, M. Guerini, · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, L. Zettlemoyer, · 2020
Cited alongside, same era.
Unifying cross-lingual semantic role labeling with heterogeneous linguistic resources,
S. Conia, A. Bacciu, R. Navigli, · 2021
Cited alongside, same era.
Point transformer,
H. Zhao, L. Jiang, J. Jia, P. H. Torr, V. Koltun, · 2021
Cited alongside, same era.
2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models,
E. B. Zaken, Y. Goldberg, S. Ravfogel, · 2022
Later among the works it cites.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks,
X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, J. Tang, · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, · 2022
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data,
C. Xu, D. Guo, N. Duan, J. McAuley, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Trappolini, L. Cosmo, L. Moschella, R. Marin, S. Melzi, E. Rodolà, · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al., · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer,
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, C. Raffel, · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation,
X. L. Li, P. Liang, · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning,
B. Lester, R. Al-Rfou, N. Constant, · 2021
Cited alongside, same era.
Do we name the languages we study? the# benderrule in lrec and acl articles,
F. Ducel, K. Fort, G. Lejeune, Y. Lepage, · 2022
Cited alongside, same era.
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al., · 2022
Cited alongside, same era.
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, G. Lample, · 2023
Closest in time.
Multimodal neural databases,
G. Trappolini, A. Santilli, E. Rodolà, A. Halevy, F. Silvestri, · 2023
Closest in time.
Musiclm: Generating music from text,
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, et al., · 2023
Closest in time.
Cycledrums: automatic drum arrangement for bass lines using cyclegan,
G. Barnabò, G. Trappolini, L. Lastilla, C. Campagnano, A. Fan, F. Petroni, F. Silvestri, · 2023
Closest in time.
Integrating item relevance in training loss for sequential recommender systems,
A. Bacciu, F. Siciliano, N. Tonellotto, F. Silvestri, · 2023
Closest in time.
A. Santilli, Camoscio: An italian instruction-tuned llama, https://github.com/teelinsan/camoscio , 2023
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, T. B. Hashimoto, Stanford alpaca: An instruction-following llama model, https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Closest in time.
A survey on efficient training of transformers,
B. Zhuang, J. Liu, Z. Pan, H. He, Y. Weng, C. Shen, · 2023
Closest in time.
W. Jiao, W. Wang, J. tse Huang, X. Wang, Z. Tu, Is chatgpt a good translator? yes with gpt-4 as the engine, 2023
2023
Closest in time.