Fetching the paper…
Reading the bibliography…
Recently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Huang, K.; Altosaar, J.; and Ranganath, R. 2019 · 1904
Earlier work this paper cites.
Cross-lingual ability of multilingual bert: An empirical study
Wang, Z.; Mayhew, S.; Roth, D.; et al. 2019 · 1912
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Celikyilmaz, A.; Clark, E.; and Gao, J. 2020 · 2006
Earlier work this paper cites.
E-BERT: A phrase and product knowledge enhanced language model for e-commerce
Zhang, D.; Yuan, Z.; Liu, Y.; Zhuang, F.; Chen, H.; and Xiong, H. 2020a · 2009
Earlier work this paper cites.
SemEval-2014 Task 4: Aspect Based Sentiment Analysis
Pontiki, M.; Galanis, D.; Pavlopoulos, J.; Papageorgiou, H.; Androutsopoulos, I.; and Manandhar, S. 2014 · 2014
Earlier work this paper cites.
Fixing weight decay regularization in adam. arXiv 2017
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
SciBERT: A Pretrained Language Model for Scientific Text
Beltagy, I.; Lo, K.; and Cohan, A. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
A dynamic product-aware learning model for e-commerce query intent understanding
Zhao, J.; Chen, H.; and Yin, D. 2019 · 2019
Earlier work this paper cites.
Jointmap: joint query intent understanding for modeling intent hierarchies in e-commerce search
Ahmadvand, A.; Kallumadi, S.; Javed, F.; and Agichtein, E. 2020 · 2020
Earlier work this paper cites.
The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service
Chen, M.; Liu, R.; Shen, L.; Yuan, S.; Zhou, J.; Wu, Y.; He, X.; and Zhou, B. 2020 · 2020
Earlier work this paper cites.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Lee, J.; Yoon, W.; Kim, S.; Kim, D.; Kim, S.; So, C. H.; and Kang, J. 2020 · 2020
Earlier work this paper cites.
The effect of natural distribution shift on question answering models
Miller, J.; Krauth, K.; Recht, B.; and Schmidt, L. 2020 · 2020
Earlier work this paper cites.
E-BERT: Efficient-Yet-Effective Entity Embeddings for BERT
Poerner, N.; Waltinger, U.; and Schütze, H. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Rastogi, A.; Zang, X.; Sunkara, S.; Gupta, R.; and Khaitan, P. 2020 · 2020
Cited alongside, same era.
Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product
Zhu, T.; Wang, Y.; Li, H.; Wu, Y.; He, X.; and Zhou, B. 2020 · 2020
Cited alongside, same era.
An end-to-end solution for named entity recognition in ecommerce search
Cheng, X.; Bowden, M.; Bhange, B. R.; Goyal, P.; Packer, T.; and Javed, F. 2021 · 2021
Cited alongside, same era.
Sustainability in e-commerce packaging: A review
Escursell, S.; Llorach-Massana, P.; and Roncero, M. B. 2021 · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; et al. 2021 · 2021
Cited alongside, same era.
Pre-training Tasks for User Intent Detection and Embedding Retrieval in E-commerce Search
Qiu, Y.; Zhao, C.; Zhang, H.; Zhuo, J.; Li, T.; Zhang, X.; Wang, S.; Xu, S.; Long, B.; and Yang, W.-Y. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Later among the works it cites.
Smith, S.; Patwary, M.; Norick, B.; LeGresley, P.; Rajbhandari, S.; Casper, J.; Liu, Z.; Prabhumoye, S.; Zerveas, G.; Korthikanti, V.; et al. 2022 · 2022
Later among the works it cites.
Galactica: A large language model for science
Taylor, R.; Kardas, M.; Cucurull, G.; Scialom, T.; Hartshorn, A.; Saravia, E.; Poulton, A.; Kerkez, V.; and Stojnic, R. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanh, V.; Webson, A.; Raffel, C.; Bach, S. H.; Sutawika, L.; Alyafeai, Z.; Chaffin, A.; Stiegler, A.; Scao, T. L.; Raja, A.; et al. 2021 · 2021
Cited alongside, same era.
Challenges and research opportunities in ecommerce search and recommendations
Tsagkias, M.; King, T. H.; Kallumadi, S.; Murdock, V.; and de Rijke, M. 2021 · 2021
Cited alongside, same era.
Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
Wang, X.; Jiang, Y.; Bach, N.; Wang, T.; Huang, Z.; Huang, F.; and Tu, K. 2021 · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Cited alongside, same era.
K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce
Xu, S.; Li, H.; Yuan, P.; Wang, Y.; Wu, Y.; He, X.; Liu, Y.; and Zhou, B. 2021 · 2021
Cited alongside, same era.
Billion-scale pre-trained e-commerce product knowledge graph model
Zhang, W.; Wong, C.-M.; Ye, G.; Wen, B.; Zhang, W.; and Chen, H. 2021 · 2021
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; Casas, D. d. L.; Hendricks, L. A.; Welbl, J.; Clark, A.; et al. 2022 · 2022
Cited alongside, same era.
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Naik, A.; Ashok, A.; Dhanasekaran, A. S.; Arunkumar, A.; Stap, D.; et al. 2022b · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Later among the works it cites.
Not All Tasks Are Born Equal: Understanding Zero-Shot Generalization
Zhou, J.; Lin, Z.; Zheng, Y.; Li, J.; and Yang, Z. 2022 · 2022
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Closest in time.
ChatLaw: Open-Source Legal Large Language Model with Integrated External Knowledge Bases
Cui, J.; Li, Z.; Yan, Y.; Chen, B.; and Yuan, L. 2023 · 2023
Closest in time.
Huang, Q.; Tao, M.; An, Z.; Zhang, C.; Jiang, C.; Chen, Z.; Wu, Z.; and Feng, Y. 2023 · 2023
Closest in time.
Penedo, G.; Malartic, Q.; Hesslow, D.; Cojocaru, R.; Cappelli, A.; Alobeidli, H.; Pannier, B.; Almazrouei, E.; and Launay, J. 2023 · 2023
Closest in time.
Large language models encode clinical knowledge
Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Wu, S.; Irsoy, O.; Lu, S.; Dabravolski, V.; Dredze, M.; Gehrmann, S.; Kambadur, P.; Rosenberg, D.; and Mann, G. 2023 · 2023
Closest in time.
FinGPT: Open-Source Financial Large Language Models
Yang, H.; Liu, X.-Y.; and Wang, C. D. 2023 · 2023
Closest in time.
XuanYuan 2.0: A Large Chinese Financial Chat Model with Hundreds of Billions Parameters
Zhang, X.; Yang, Q.; and Xu, D. 2023 · 2023
Closest in time.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Closest in time.