Fetching the paper…
Reading the bibliography…
Interacting with human via high-quality multi-turn dialogues is a key feature of large language models (LLMs).
The proposed uscf rating system, its development, theory, and applications
Elo, A. E · 1967
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
A survey of available corpora for building data-driven dialogue systems
Serban, I. V., Lowe, R., Henderson, P., Charlin, L., and Pineau, J · 2015
Earlier work this paper cites.
Liu, C.-W., Lowe, R., Serban, I. V., Noseworthy, M., Charlin, L., and Pineau, J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Dinan, E., Roller, S., Shuster, K., Fan, A., Auli, M., and Weston, J · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Zhang, S., Dinan, E., Urbanek, J., Szlam, A., Kiela, D., and Weston, J · 2018
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
Fabbri, A. R., Li, I., She, T., Li, S., and Radev, D. R · 2019
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Li, M., Weston, J., and Roller, S · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Reddy, S., Chen, D., and Manning, C. D · 2019
Cited alongside, same era.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., et al · 2020
Cited alongside, same era.
Mutual: A dataset for multi-turn dialogue reasoning
Cui, L., Wu, Y., Liu, S., Zhang, Y., and Zhou, M · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Closest in time.
Gptscore: Evaluate as you desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Closest in time.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., Liu, J., Lv, C., Zhang, Y., Lei, J., Fu, Y., Sun, M., and He, J · 2023
Closest in time.
Is chatgpt a good translator? a preliminary study
Jiao, W., Wang, W., Huang, J.-t., Wang, X., and Tu, Z · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L · 2021
Cited alongside, same era.
Doğruöz, A. S. and Skantze, G · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al · 2022
Cited alongside, same era.
Qwen technical report, 2023
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Cited alongside, same era.
Baichuan 2: Open large-scale language models
Baichuan · 2023
Cited alongside, same era.
Emergent autonomous scientific research capabilities of large language models
Boiko, D. A., MacKnight, R., and Gomes, G · 2023
Cited alongside, same era.
Chemcrow: Augmenting large-language models with chemistry tools
Bran, A. M., Cox, S., White, A. D., and Schwaller, P · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023a
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G
Cited in the paper.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Closest in time.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Closest in time.
Is chatgpt a good nlg evaluator? a preliminary study
Wang, J., Liang, Y., Meng, F., Shi, H., Li, Z., Xu, J., Qu, J., and Zhou, J · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions, 2023
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Closest in time.
A dataset for document grounded conversations
Zhou, K., Prabhumoye, S., and Black, A. W · 2023
Closest in time.