Fetching the paper…
Reading the bibliography…
Large language models (LLMs) like Llama, Baichuan and Bloom models show remarkable ability with instruction fine-tuning in many natural language tasks.
Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020, 33: 1877-1901
1901
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, et al. Bleu: a method for automatic evaluation of machine translation[C]//Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 2002: 311-318
2002
Earlier work this paper cites.
Lin C Y. Rouge: A package for automatic evaluation of summaries[C]//Text summarization branches out. 2004: 74-81
2004
Earlier work this paper cites.
2015
Earlier work this paper cites.
Liu P J, Manning C D. Get to the point: Summarization with pointer-generator networks[C]//Proc. Conf. Annu. Meeting Assoc. Comput. Linguistics (ACL). 2017: 1073-1083
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Kenton J D M W C, Toutanova L K. Bert: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of naacL-HLT. 2019, 1: 2
2019
Earlier work this paper cites.
Gliwa B, Mochol I, Biesek M, et al. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization[J]. ar**v preprint ar**v:1911.12237, 2019
2019
Earlier work this paper cites.
Liu Y, Lapata M. Text summarization with pretrained encoders[J]. ar**v preprint ar**v:1908.08345, 2019
2019
Earlier work this paper cites.
Zhang B, Sennrich R. Root mean square layer normalization[J]. Advances in Neural Information Processing Systems, 2019, 32
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Feng X, Feng X, Qin B, et al. Dialogue discourse-aware graph convolutional networks for abstractive meeting summarization[J]. ar**v preprint ar**v:2012.03502, 2020, 17
2020
Earlier work this paper cites.
Lewis M, Liu Y, Goyal N, et al. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020: 7871-7880
2020
Earlier work this paper cites.
Shazeer N. Glu variants improve transformer[J]. arXiv preprint arXiv:2002.05202, 2020
2020
Cited alongside, same era.
Zou Y, Zhao L, Kang Y, et al. Topic-oriented spoken dialogue summarization for customer service with saliency-aware topic modeling[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(16): 14665-14673
2021
Cited alongside, same era.
Lin H, Ma L, Zhu J, et al. CSDS: A fine-grained Chinese dataset for customer service dialogue summarization[J]. ar**v preprint ar**v:2108.13139, 2021
2021
Cited alongside, same era.
Feng X, Feng X, Qin L, et al. Language model as an annotator: Exploring DialoGPT for dialogue summarization[J]. ar**v preprint ar**v:2105.12544, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2023
Later among the works it cites.
Touvron H, Martin L, Stone K, et al. Llama 2: Open foundation and fine-tuned chat models[J]. ar**v preprint ar**v:2307.09288, 2023
2023
Later among the works it cites.
Yang A, **ao B, Wang B, et al. Baichuan 2: Open large-scale language models[J]. ar**v preprint ar**v:2309.10305, 2023
2023
Later among the works it cites.
Jain N, Chiang P, Wen Y, et al. NEFTune: Noisy Embeddings Improve Instruction Finetuning[J]. ar**v preprint ar**v:2310.05914, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Du Z, Qian Y, Liu X, et al. Glm: General language model pretraining with autoregressive blank infilling[J]. ar**v preprint ar**v:2103.10360, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Hu E, Shen Y, Wallis P, et al. LoRA: Low-Rank Adaptation of Large Language Models[J]. 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Liang X, Bian C, Wu S, et al. Towards Modeling Role-Aware Centrality for Dialogue Summarization[C]//Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing. 2022: 43-50
2022
Cited alongside, same era.
Lin H, Zhu J, **ang L, et al. Other roles matter! enhancing role-oriented dialogue summarization via role interactions[J]. ar**v preprint ar**v:2205.13190, 2022
2022
Cited alongside, same era.
Zhong M, Liu Y, Xu Y, et al. Dialoglm: Pre-trained model for long dialogue understanding and summarization[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2022, 36(10): 11765-11773
2022
Cited alongside, same era.
2023
Later among the works it cites.
https://crfm.stanford.edu/2023/03/13/alpaca.html
2023
Later among the works it cites.
Liu H, Li C, Wu Q, et al. Visual instruction tuning[J]. arXiv preprint arXiv:2304.08485, 2023
2023
Later among the works it cites.
Liu X, Zheng Y, Du Z, et al. GPT understands, too[J]. AI Open, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Hugging Face. Enhance model’s performances using NEFTune. https://huggingface.co/docs/trl/main/en/sft_trainer, 2023
2023
Later among the works it cites.
Su J, Ahmed M, Lu Y, et al. Roformer: Enhanced transformer with rotary position embedding[J]. Neurocomputing, 2024, 568: 127063
2024
Closest in time.