Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models (LLMs) and Reinforcement Learning (RL) have shown significant promise in decision-making tasks.
Markov decision processes
Puterman, M. L · 1990
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W. B. and Stone, P · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Agent-agnostic human-in-the-loop reinforcement learning
Abel, D., Salvatier, J., Stuhlmüller, A., and Evans, O · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L · 2017
Earlier work this paper cites.
A berkeley view of systems challenges for ai
Stoica, I., Song, D., Popa, R. A., Patterson, D., Mahoney, M. W., Katz, R., Joseph, A. D., Jordan, M., Hellerstein, J. M., Gonzalez, J. E., et al · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Hierarchical imitation and reinforcement learning
Le, H., Jiang, N., Agarwal, A., Dudík, M., Yue, Y., and Daumé III, H · 2018
Earlier work this paper cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Earlier work this paper cites.
Learning ambidextrous robot grasping policies
Mahler, J., Matl, M., Satish, V., Danielczuk, M., DeRose, B., McKinley, S., and Goldberg, K · 2019
Earlier work this paper cites.
Grid2op-a testbed platform to model sequential decision making in power systems
Donnot, B · 2020
Earlier work this paper cites.
Learning to run a power network challenge for training topology controllers
Marot, A., Donnot, B., Romero, C., Donon, B., Lerousseau, M., Veyrin-Forrer, L., and Guyon, I · 2020
Cited alongside, same era.
A review on reinforcement learning: Introduction and applications in industrial process control
Nian, R., Liu, J., and Huang, B · 2020
Cited alongside, same era.
Keep calm and explore: Language models for action generation in text-based games
Yao, S., Rao, R., Hausknecht, M., and Narasimhan, K · 2020
Cited alongside, same era.
Large language model as a policy teacher for training reinforcement learning agents
Zhou, Z., Hu, B., Zhao, C., Zhang, P., and Liu, B · 2020
Cited alongside, same era.
Learning to run a power network challenge: a retrospective analysis
Marot, A., Donnot, B., Dulac-Arnold, G., Kelly, A., O’Sullivan, A., Viebahn, J., Awad, M., Guyon, I., Panciatici, P., and Romero, C · 2021
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Later among the works it cites.
Powrl: A reinforcement learning framework for robust management of power networks
Chauhan, A., Baranwal, M., and Basumatary, A · 2023
Later among the works it cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Later among the works it cites.
Lift: Unsupervised reinforcement learning with foundation models as teachers
Nam, T., Lee, J., Zhang, J., Hwang, S. J., Lim, J. J., and Pertsch, K · 2023
Later among the works it cites.
Soft dagger: Sample-efficient imitation learning for control of soft robots
Nazeer, M. S., Laschi, C., and Falotico, E · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Winning the l2rpn challenge: Power grid management via semi-markov afterstate actor-critic
Yoon, D., Hong, S., Lee, B.-J., and Kim, K.-E · 2021
Cited alongside, same era.
A survey of inverse reinforcement learning
Adams, S., Cody, T., and Beling, P. A · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Power grid congestion management via topology optimization with alphazero
Dorfer, M., Fuxjäger, A. R., Kozak, K., Blies, P. M., and Wasserer, M · 2022
Cited alongside, same era.
Scaling up multi-task robotic reinforcement learning
Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K · 2022
Cited alongside, same era.
Learning to run a power network with trust
Marot, A., Donnot, B., Chaouache, K., Kelly, A., Huang, Q., Hossain, R.-R., and Cremer, J. L · 2022
Cited alongside, same era.
Later among the works it cites.
Lgts: Dynamic task sampling using llm-generated sub-goals for reinforcement learning agents
Shukla, Y., Gao, W., Sarathy, V., Velasquez, A., Wright, R., and Sinapov, J · 2023
Later among the works it cites.
Tpe: Towards better compositional reasoning over conceptual tools with multi-persona collaboration
Wang, H., Wang, H., Wang, L., Hu, M., Wang, R., Xue, B., Lu, H., Mi, F., and Wong, K.-F · 2023
Later among the works it cites.
Zeng, S., Li, C., Garcia, A., and Hong, M · 2023
Later among the works it cites.
Serverlessllm: Low-latency serverless inference for large language models
Fu, Y., Xue, L., Huang, Y., Brabete, A.-O., Ustiugov, D., Patel, Y., and Mai, L · 2024
Later among the works it cites.
Applications, challenges, and future directions of human-in-the-loop learning
Kumar, S., Datta, S., Singh, V., Datta, D., Singh, S. K., and Sharma, R · 2024
Later among the works it cites.
Rl-gpt: Integrating reinforcement learning and code-as-policy
Liu, S., Yuan, H., Hu, M., Li, Y., Chen, Y., Liu, S., Lu, Z., and Jia, J · 2024
Later among the works it cites.
Chameleon: Plug-and-play compositional reasoning with large language models
Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K.-W., Wu, Y. N., Zhu, S.-C., and Gao, J · 2024
Later among the works it cites.
Tan, W., Zhang, W., Liu, S., Zheng, L., Wang, X., and An, B · 2024
Later among the works it cites.
Dart-llm: Dependency-aware multi-robot task decomposition and execution using large language models
Wang, Y., Xiao, R., Kasahara, J. Y. L., Yajima, R., Nagatani, K., Yamashita, A., and Asama, H · 2024
Later among the works it cites.
Reinforcing llm agents via policy optimization with action decomposition
Wen, M., Wan, Z., Wang, J., Zhang, W., and Wen, Y · 2024
Later among the works it cites.
How can llm guide rl? a value-based approach
Zhang, S., Zheng, S., Ke, S., Liu, Z., Jin, W., Yuan, J., Yang, Y., Yang, H., and Wang, Z · 2024
Later among the works it cites.