Fetching the paper…
Reading the bibliography…
This paper presents an innovative framework that integrates Large Language Models (LLMs) with an external Thinker module to enhance the reasoning capabilities of LLM-based agents.
Dual processes in reasoning?
Wason, P. C. and Evans, J. S. B · 1974
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Earlier work this paper cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Earlier work this paper cites.
Application of deep reinforcement learning in werewolf game agents
Wang, T. and Kaneko, T · 2018
Earlier work this paper cites.
A study of ai agent commitment in one night ultimate werewolf with human players
Eger, M. and Martens, C · 2019
Earlier work this paper cites.
Finding friend and foe in multi-agent games
Serrino, J., Kleiman-Weiner, M., Parkes, D. C., and Tenenbaum, J · 2019
Earlier work this paper cites.
Shallow-fusion end-to-end contextual biasing
Zhao, D., Sainath, T. N., Rybach, D., Rondon, P., Bhatia, D., Li, B., and Pang, R · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Towards playing full moba games with deep reinforcement learning
Ye, D., Chen, G., Zhang, W., Chen, S., Yuan, B., Liu, B., Chen, J., Liu, Z., Qiu, F., Yu, H., et al · 2020
Earlier work this paper cites.
Rlupus: Cooperation through emergent communication in the werewolf social deduction game
Brandizzi, N., Grossi, D., and Iocchi, L · 2021
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al · 2022
Cited alongside, same era.
Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Gao, Z., Zhang, S., McLoughlin, I., and Yan, Z · 2022
Cited alongside, same era.
Towards reasoning in large language models: A survey
Huang, J. and Chang, K. C.-C · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I · 2022
Cited alongside, same era.
Chatdb: Augmenting llms with databases as their symbolic memory
Hu, C., Fu, J., Du, C., Luo, S., Zhao, J., and Zhao, H · 2023
Later among the works it cites.
Agentsims: An open-source sandbox for large language model evaluation
Lin, J., Zhao, H., Zhang, A., Wu, Y., Ping, H., and Chen, Q · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A novel weighted ensemble learning based agent for the werewolf game
Khan, M. and Aranha, C · 2022
Cited alongside, same era.
Hidden agenda: a social deduction game with diverse learned equilibria
Kopparapu, K., Duéñez-Guzmán, E. A., Matyas, J., Vezhnevets, A. S., Agapiou, J. P., McKee, K. R., Everett, R., Marecki, J., Leibo, J. Z., and Graepel, T · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Cited alongside, same era.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Wu, T., Terry, M., and Cai, C. J · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Cited alongside, same era.
Communicative agents for software development, 2023
Qian, C., Cong, X., Liu, W., Yang, C., Chen, W., Su, Y., Dang, Y., Li, J., Xu, J., Li, D., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Playing the werewolf game with artificial intelligence for language understanding
Shibata, H., Miki, S., and Nakamura, Y · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S · 2023
Later among the works it cites.
Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems
Stechly, K., Marquez, M., and Kambhampati, S · 2023
Later among the works it cites.
Can large language models really improve by self-critiquing their own plans?
Valmeekam, K., Marquez, M., and Kambhampati, S · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang, Z., Li, L., Wang, J., Lin, K., Azarnasab, E., Ahmed, F., Liu, Z., Liu, C., Zeng, M., and Wang, L · 2023
Later among the works it cites.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., and Wang, Y · 2023
Later among the works it cites.
Calypso: Llms as dungeon master’s assistants
Zhu, A., Martin, L., Head, A., and Callison-Burch, C · 2023
Later among the works it cites.