Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities in solving various tasks, yet they often struggle with comprehensively addressing complex and vague problems.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Open problems in cooperative ai
Dafoe, A., Hughes, E., Bachrach, Y., Collins, T., McKee, K. R., Leibo, J. Z., Larson, K., and Graepel, T. (2020) · 2012
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Emergent translation in multi-agent communication
Lee, J., Cho, K., Weston, J., and Kiela, D. (2018) · 2018
Earlier work this paper cites.
Emergent linguistic phenomena in multi-agent communication games
Graesser, L., Cho, K., and Kiela, D. (2020) · 2020
Earlier work this paper cites.
Multi-agent communication meets natural language: Synergies between functional and structural language learning
Lazaridou, A., Potapenko, A., and Tieleman, O. (2020) · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021) · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021) · 2021
Earlier work this paper cites.
Cooperative ai: machines must learn to find common ground
Dafoe, A., Bachrach, Y., Hadfield, G., Horvitz, E., Larson, K., and Graepel, T. (2021) · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. (2021) · 2021
Earlier work this paper cites.
Langchain, october 2022
Chase, H. (2022) · 2022
Earlier work this paper cites.
Negotiation and honesty in artificial intelligence methods for the board game of diplomacy
Kramár, J., Eccles, T., Gemp, I., Tacchetti, A., McKee, K. R., Malinowski, M., Graepel, T., and Bachrach, Y. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022) · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022) · 2022
Cited alongside, same era.
Large language models and the perils of their hallucinations
Camel: Communicative agents for" mind" exploration of large scale language model society
Li, G., Hammoud, H. A. A. K., Itani, H., Khizbullin, D., and Ghanem, B. (2023) · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S. (2023) · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
OpenAI (2023) · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Azamfirei, R., Kudchadkar, S. R., and Fackler, J. (2023) · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Chiang, C.-H. and Lee, H.-y. (2023) · 2023
Cited alongside, same era.
Lm vs lm: Detecting factual errors via cross examination
Cohen, R., Hamri, M., Geva, M., and Globerson, A. (2023) · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023) · 2023
Cited alongside, same era.
Human-like summarization evaluation with chatgpt
Gao, M., Ruan, J., Sun, R., Yin, X., Yang, S., and Wan, X. (2023) · 2023
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A. (2023) · 2023
Cited alongside, same era.
Liu, J., Xia, C. S., Wang, Y., and Zhang, L. (2023a)
Cited in the paper.
Qian, C., Cong, X., Yang, C., Chen, W., Su, Y., Xu, J., Liu, Z., and Sun, M. (2023) · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al. (2023) · 2023
Later among the works it cites.
Are large language models good evaluators for abstractive summarization?
Shen, C., Cheng, L., You, Y., and Bing, L. (2023) · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Later among the works it cites.
Large language models are diverse role-players for summarization evaluation
Wu, N., Gong, M., Shou, L., Liang, S., and Jiang, D. (2023) · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Later among the works it cites.
Llm harmony: Multi-agent communication for problem solving
Rasal, S. (2024) · 2024
Closest in time.