Fetching the paper…
Reading the bibliography…
Achieving the effective design and improvement of reward functions in reinforcement learning (RL) tasks with complex custom environments and multiple requirements presents considerable challenges.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International conference on machine learning . PMLR, 2018, pp. 1587–1596
2018
Earlier work this paper cites.
C. F. Hayes, R. Rădulescu, E. Bargiacchi, J. Källström, M. Macfarlane, M. Reymond, T. Verstraeten, L. M. Zintgraf, R. Dazeley, F. Heintz et al. , “A practical guide to multi-objective reinforcement learning and planning,” Autonomous Agents and Multi-Agent Systems , vol. 36, no. 1, p. 26, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Zhang, N. Desai, J. Bae, J. Lorraine, and J. Ba, “Using large language models for hyperparameter optimization,” in NeurIPS 2023 Foundation Models for Decision Making Workshop , 2023
2023
Earlier work this paper cites.
C. Wang, X. Liu, and A. H. Awadallah, “Cost-effective hyperparameter optimization for large language model generation inference,” in International Conference on Automated Machine Learning . PMLR, 2023, pp. 21–1
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar, “Eureka: Human-level reward design via coding large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Wu, S. Y. Min, S. Prabhumoye, Y. Bisk, R. R. Salakhutdinov, A. Azaria, T. M. Mitchell, and Y. Li, “Spring: Studying papers and reasoning to play games,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
H. Li, X. Yang, Z. Wang, X. Zhu, J. Zhou, Y. Qiao, X. Wang, H. Li, L. Lu, and J. Dai, “Auto mc-reward: Automated dense reward design with large language models for minecraft,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
2024
Closest in time.
2024
Closest in time.
Z. Zhang, J. Xu, G. Xie, J. Wang, Z. Han, and Y. Ren, “Environment- and energy-aware ama-assisted data collection for the information updating networks,” IEEE Internet of Things Journal , vol. 11, no. 15, pp. 26 406–26 418, 2024
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Z. Mandi, S. Jain, and S. Song, “Roco: Dialectic multi-robot collaboration with large language models,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 286–299
2024
Cited alongside, same era.
E. Triantafyllidis, F. Christianos, and Z. Li, “Intrinsic language-guided exploration for complex long-horizon robotic manipulation tasks,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 7493–7500
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu, “Text2reward: Reward shaping with language models for reinforcement learning,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.