Fetching the paper…
Reading the bibliography…
Building embodied agents on integrating Large Language Models (LLMs) and Reinforcement Learning (RL) have revolutionized human-AI interaction: researchers can now leverage language instructions to plan decision-making for open-ended tasks.
The string-to-string correction problem
Wagner, R. A. and Fischer, M. J · 1974
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Grounded language learning in a simulated 3d world
Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Guided goal generation for hindsight multi-goal reinforcement learning
Bai, C., Liu, P., Zhao, W., and Tang, X · 2019
Earlier work this paper cites.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Perolat, J., Jaderberg, M., and Graepel, T · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning, 2019
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Earlier work this paper cites.
Curriculum-guided hindsight experience replay
Fang, M., Zhou, T., Du, Y., Han, L., and Zhang, Z · 2019
Earlier work this paper cites.
Exploration via hindsight goal generation
Ren, Z., Dong, K., Zhou, Y., Liu, Q., and Peng, J · 2019
Cited alongside, same era.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards, 2019
Trott, A., Zheng, S., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Curriculum for reinforcement learning
Weng, L · 2020
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
Khandelwal, A., Weihs, L., Mottaghi, R., and Kembhavi, A · 2022
Later among the works it cites.
Goal-conditioned reinforcement learning: Problems and solutions
Liu, M., Zhu, M., and Zhang, W · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A · 2022
Later among the works it cites.
Magnetic field-based reward shaping for goal-conditioned reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2021
Cited alongside, same era.
Battle royale: First-person shooter game
Gautam, A., Jain, H., Senger, A., and Dhand, G · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Cited alongside, same era.
Open-ended learning leads to generally capable agents
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., et al · 2021
Cited alongside, same era.
Noveld: A simple yet effective exploration criterion
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y · 2021
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al · 2021
Cited alongside, same era.
Ding, H., Tang, Y., Wu, Q., Wang, B., Chen, C., and Wang, Z · 2023
Closest in time.
Guiding pretraining in reinforcement learning with large language models
Du, Y., Watkins, O., Wang, Z., Colas, C., Darrell, T., Abbeel, P., Gupta, A., and Andreas, J · 2023
Closest in time.
Enabling intelligent interactions between an agent and an llm: A reinforcement learning approach, 2023
Hu, B., Zhao, C., Zhang, P., Zhou, Z., Yang, Y., Xu, Z., and Liu, B · 2023
Closest in time.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Fei-Fei, L · 2023
Closest in time.
Vima: Robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2023
Closest in time.
Anti-exploration by random network distillation
Nikulin, A., Kurenkov, V., Tarasov, D., and Kolesnikov, S · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Closest in time.