Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have recently been used for sequential decision making in interactive environments.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Earlier work this paper cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., and Rocktäschel, T · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Codet: Code generation with generated tests
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., and Chen, W · 2022
Earlier work this paper cites.
RLPrompt: Optimizing discrete text prompts with reinforcement learning
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E., and Hu, Z · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ichter, B., Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., Kalashnikov, D., Levine, S., Lu, Y., Parada, C., Rao, K., Sermanet, P., Toshev, A. T., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Yan, M., Brown, N., Ahn, M., Cortes, O., Sievers, N., Tan, C., Xu, S., Reyes, D., Rettinghouse, J., Quiambao, J., Pastor, P., Luu, L., Lee, K.-H., Kuang, Y., Jesmonth, S., Jeffrey, K., Ruano, R. J., Hsu, J., Gopalakrishnan, K., David, B., Zeng, A., and Fu, C. K · 2022
Earlier work this paper cites.
Scienceworld: Is your agent smarter than a 5th grader?
Wang, R., Jansen, P., Côté, M.-A., and Ammanabrolu, P · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2022
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T · 2023
Cited alongside, same era.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Cited alongside, same era.
Clin: A continually learning language agent for rapid task adaptation and generalization
Majumder, B. P., Mishra, B. D., Jansen, P., Tafjord, O., Tandon, N., Zhang, L., Callison-Burch, C., and Clark, P · 2023
Later among the works it cites.
Selective perception: Learning concise state descriptions for language model actors
Nottingham, K., Razeghi, Y., Kim, K., Lanier, J., Baldi, P., Fox, R., and Singh, S · 2023
Later among the works it cites.
Adapt: As-needed decomposition and planning with language models
Prasad, A., Koller, A., Hartmann, M., Clark, P., Sabharwal, A., Bansal, M., and Khot, T · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Later among the works it cites.
Tempera: Test-time prompt editing via reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liang, Y., Wu, C., Song, T., Wu, W., Xia, Y., Liu, Y., Ou, Y., Lu, S., Ji, L., Mao, S., et al · 2023
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I
Cited in the paper.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al
Cited in the paper.
Do embodied agents dream of pixelated sheep? embodied decision making using language guided world modelling
Nottingham, K., Ammanabrolu, P., Suhr, A., Choi, Y., Hajishirzi, H., Singh, S., and Fox, R
Cited in the paper.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A
Cited in the paper.
Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents
Wang, Z., Cai, S., Chen, G., Liu, A., Ma, X., and Liang, Y
Cited in the paper.
Zhang, T., Wang, X., Zhou, D., Schuurmans, D., and Gonzalez, J. E · 2023
Later among the works it cites.
Expel: Llm agents are experiential learners
Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., and Huang, G · 2023
Later among the works it cites.