Fetching the paper…
Reading the bibliography…
Hybrid Group Relative Policy Optimization (Hybrid GRPO) is a reinforcement learning framework that extends Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO) by incorporating empirical multi-sample action evaluation while preserving the stability of value function-based learning.
Ziebart, B. D. (2008). Maximum entropy reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS)
2008
Earlier work this paper cites.
2015
Earlier work this paper cites.
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (ICML)
2016
Earlier work this paper cites.
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., and Lanctot, M. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning (ICML)
2018
Cited alongside, same era.
Sutton, R. S., and Barto, A. G. (2018). Reinforcement learning: An introduction. MIT Press
2018
Cited alongside, same era.
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D. (2018). Distributed prioritized experience replay. In International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
Tesla, Inc. (2023). Full Self-Driving (FSD) software overview. Tesla AI Research
2023
Later among the works it cites.
DeepSeek. (2025). Group Relative Policy Optimization (GRPO). Technical report
2025
Closest in time.
Sane, S., Stable-Baselines3 Development Team. (2025). Group Relative Policy Optimization (GRPO) GitHub Repository. Available at: https://github.com/Soham4001A/stable-baselines3-contrib . Accessed: January 27, 2025
2025
Closest in time.
Sane, S. (2025). Hybrid GRPO: Reinforcement Learning Framework GitHub Repository. Available at: https://github.com/Soham4001A/RL_Tracking . Accessed: January 27, 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.