Can large language models be an alternative to human evaluations?
Original
Chiang, C.-H. and Lee, H.-y. (2023) · 2023
Later among the works it cites.
Lm vs lm: Detecting factual errors via cross examination
Original
Cohen, R., Hamri, M., Geva, M., and Globerson, A. (2023) · 2023
Later among the works it cites.
Improving factuality and reasoning in language models through multiagent debate
Original
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023) · 2023
Later among the works it cites.
Human-like summarization evaluation with chatgpt
Original
Gao, M., Ruan, J., Sun, R., Yin, X., Yang, S., and Wan, X. (2023) · 2023
Later among the works it cites.
Camel: Communicative agents for" mind" exploration of large scale language model society
Original
Li, G., Hammoud, H. A. A. K., Itani, H., Khizbullin, D., and Ghanem, B. (2023) · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Original
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S. (2023) · 2023
Later among the works it cites.
Chameleon: Plug-and-play compositional reasoning with large language models
Original
Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K.-W., Wu, Y. N., Zhu, S.-C., and Gao, J. (2023) · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
OpenAI (2023) · 2023
Later among the works it cites.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D. (2023) · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Original
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E. (2023) · 2023
Later among the works it cites.
Communicative agents for software development
Original
Qian, C., Cong, X., Yang, C., Chen, W., Su, Y., Xu, J., Liu, Z., and Sun, M. (2023) · 2023
Later among the works it cites.
Beyond segmentation: Road network generation with multi-modal llms
Original
Rasal, S. and Boddhu, S. K. (2023) · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Original
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023) · 2023
Later among the works it cites.
Minding language models’(lack of) theory of mind: A plug-and-play multi-character belief tracker
Original
Sclar, M., Kumar, S., West, P., Suhr, A., Choi, Y., and Tsvetkov, Y. (2023) · 2023
Later among the works it cites.
Are large language models good evaluators for abstractive summarization?
Original
Shen, C., Cheng, L., You, Y., and Bing, L. (2023) · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Original
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Later among the works it cites.
Large language models are diverse role-players for summarization evaluation
Original
Wu, N., Gong, M., Shou, L., Liang, S., and Jiang, D. (2023) · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Original
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Later among the works it cites.
I cast detect thoughts: Learning to converse and guide with intents and theory-of-mind in dungeons and dragons
Original
Zhou, P., Zhu, A., Hu, J., Pujara, J., Ren, X., Callison-Burch, C., Choi, Y., and Ammanabrolu, P. (2023) · 2023
Later among the works it cites.