Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has demonstrated impressive performance in decision-making tasks like embodied control, autonomous driving and financial trading.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM journal on control and optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming , ser. Wiley Series in Probability and Statistics. Wiley, 1994
1994
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-dynamic programming , ser. Optimization and neural computation series. Athena Scientific, 1996
1996
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, “Reinforcement learning: An introduction,” IEEE Trans. Neural Networks , vol. 9, no. 5, pp. 1054–1054, 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in NIPS , 1999
1999
Earlier work this paper cites.
R. Munos, “Error bounds for approximate value iteration,” in AAAI , 2005
2005
Earlier work this paper cites.
L. Kocsis and C. Szepesvári, “Discounted ucb,” in PASCAL Challenges Workshop , 2006
2006
Earlier work this paper cites.
L.-J. Ji, Z. Zhang, and T. Guo, “To buy or to sell: Cultural differences in stock market decisions based on price trends,” Journal of Behavioral Decision Making , vol. 21, no. 4, pp. 399–413, 2008
2008
Earlier work this paper cites.
A. Garivier and E. Moulines, “On upper-confidence bound policies for non-stationary bandit problems,” 2008
2008
Earlier work this paper cites.
M. E. Taylor, H. B. Suay, and S. Chernova, “Integrating reinforcement learning with human demonstrations of varying ability,” in AAMAS , 2011
2011
Earlier work this paper cites.
H. Akiyama and T. Nakashima, “HELIOS2012: robocup 2012 soccer simulation 2d league champion,” in ARIS , 2012
2012
Earlier work this paper cites.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in NeurIPS , 2014
2014
Earlier work this paper cites.
Y. Gur, A. Zeevi, and O. Besbes, “Stochastic multi-armed-bandit problem with non-stationary rewards,” in NeurIPS , 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
M. Petrik and R. Luss, “Interpretable policies for dynamic product recommendations,” in UAI , 2016
2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” 2016
2016
Earlier work this paper cites.
M. Hausknecht, P. Mupparaju, S. Subramanian, S. Kalyanakrishnan, and P. Stone, “Half field offense: An environment for multiagent learning and ad hoc teamwork,” in AAMAS ALA Workshop , 2016
2016
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017
2017
Earlier work this paper cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in ICML , 2017
2017
Earlier work this paper cites.
L. Dong, X. Zhong, C. Sun, and H. He, “Event-triggered adaptive dynamic programming for continuous-time systems with control constraints,” IEEE Transactions on Neural Networks and Learning Systems , vol. 28, no. 8, pp. 1941–1952, 2017
2017
Earlier work this paper cites.
N. Cesa-Bianchi, C. Gentile, G. Neu, and G. Lugosi, “Boltzmann exploration done right,” in NeurIPS , 2017
2017
Earlier work this paper cites.
“Sina finance,” https://finance.sina.com.cn/stock/ , stock code: XSHE: 000025, from 10/1/2017 to 26/5/2017
2017
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in ICML , 2017
2017
Earlier work this paper cites.
X. Zhang, G. Liu, C. Yang, and J. Wu, “Research on air confrontation maneuver decision-making method based on reinforcement learning,” Electronics , vol. 7, no. 11, p. 279, 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in ICML , 2018
2018
Earlier work this paper cites.
K. Lee, S. Choi, and S. Oh, “Sparse markov decision processes with causal sparse tsallis entropy regularization for reinforcement learning,” IEEE Robotics and Automation Letters , vol. 3, no. 2, pp. 1466–1473, 2018
2018
Earlier work this paper cites.
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” in AAAI , 2018
2018
Cited alongside, same era.
A. Hill, A. Raffin, M. Ernestus, A. Gleave, A. Kanervisto, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu, “Stable baselines,” 2018
2018
Cited alongside, same era.
T. L. Meng and M. Khushi, “Reinforcement learning in financial markets,” Data , vol. 4, no. 3, pp. 110–110, 2019
2019
Cited alongside, same era.
R. Noothigattu, D. Bouneffouf, N. Mattei, R. Chandra, P. Madan, K. R. Varshney, M. Campbell, M. Singh, and F. Rossi, “Teaching AI agents ethical values using reinforcement learning and policy orchestration,” in IJCAI , 2019
2019
Cited alongside, same era.
E. Kim, H. Choi, H. Kim, J. Na, and H. Lee, “Optimal resource allocation considering non-uniform spatial traffic distribution in ultra-dense networks: A multi-agent reinforcement learning approach,” IEEE Access , vol. 10, pp. 20 455–20 464, 2022
2022
Closest in time.
C. Liu, L. Huang, and Z. Dong, “A two-stage approach of joint route planning and resource allocation for multiple uavs in unmanned logistics distribution,” IEEE Access , vol. 10, pp. 113 888–113 901, 2022
2022
Closest in time.
S. Huang and S. Ontañón, “A closer look at invalid action masking in policy gradient algorithms,” in FLAIRS , 2022
2022
Closest in time.
T. Mu, G. Theocharous, D. Arbour, and E. Brunskill, “Constraint sampling reinforcement learning: Incorporating expertise for faster learning,” in AAAI , 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Balakrishnan, D. Bouneffouf, N. Mattei, and F. Rossi, “Using multi-armed bandits to learn ethical priorities for online AI systems,” IBM Journal of Research and Development , vol. 63, no. 4/5, pp. 1–1, 2019
2019
Cited alongside, same era.
W. Yang, X. Li, and Z. Zhang, “A regularized approach to sparse optimal policy in reinforcement learning,” in NeurIPS , 2019
2019
Cited alongside, same era.
A. Trott, S. Zheng, C. Xiong, and R. Socher, “Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards,” in NeurIPS , 2019
2019
Cited alongside, same era.
M. Geist, B. Scherrer, and O. Pietquin, “A theory of regularized markov decision processes,” in ICML , 2019
2019
Cited alongside, same era.
S. Jiang, J. Pang, and Y. Yu, “Offline imitation learning with a misspecified simulator,” in NeurIPS , 2020
2020
Cited alongside, same era.
L. Li, L. Sun, C. Weng, C. Huo, and W. Ren, “Spending money wisely: Online electronic coupon allocation based on real-time user intent detection,” in CIKM , 2020
2020
Cited alongside, same era.
A. Brim, “Deep reinforcement learning pairs trading with a double deep q-network,” in CCWC , 2020
2020
Cited alongside, same era.
M. Habibullah, M. A. M. Islam, N. B. Alam, and F. Ahmed, “Player performance profiling for penalty shootouts in football using video analysis,” in ICCA , 2020
2020
Cited alongside, same era.
A. Srivastava and S. M. Salapaka, “Parameterized mdps and reinforcement learning problems - A maximum entropy principle-based framework,” IEEE Transactions on Cybernetics , vol. 52, no. 9, pp. 9339–9351, 2022
2022
Closest in time.
F. Ding and Y. Xue, “X-MEN: guaranteed xor-maximum entropy constrained inverse reinforcement learning,” in UAI , 2022
2022
Closest in time.
D. Rengarajan, G. Vaidya, A. Sarvesh, D. M. Kalathil, and S. Shakkottai, “Reinforcement learning with sparse rewards using guidance from offline demonstration,” in ICLR , 2022
2022
Closest in time.
Y. Guo, Q. Wu, and H. Lee, “Learning action translator for meta reinforcement learning on sparse-reward tasks,” in AAAI , 2022
2022
Closest in time.
S. Chakraborty, A. S. Bedi, A. Koppel, P. Tokekar, and D. Manocha, “Dealing with sparse rewards in continuous control robotics via heavy-tailed policies,” 2022
2022
Closest in time.
R. Devidze, P. Kamalaruban, and A. Singla, “Exploration-guided reward shaping for reinforcement learning under sparse rewards,” in NeurIPS , 2022
2022
Closest in time.
P. Ladosz, E. Ben-Iwhiwhu, J. Dick, N. Ketz, S. Kolouri, J. L. Krichmar, P. K. Pilly, and A. Soltoggio, “Deep reinforcement learning with modulated hebbian plus q-network architecture,” IEEE Transactions on Neural Networks and Learning Systems , vol. 33, no. 5, pp. 2045–2056, 2022
2022
Closest in time.
A. H. Tan, F. P. Bejarano, Y. Zhu, R. Ren, and G. Nejat, “Deep reinforcement learning for decentralized multi-robot exploration with macro actions,” IEEE Robotics and Automation Letters , vol. 8, no. 1, pp. 272–279, 2023
2023
Closest in time.
J. Pang, S. Yang, X. Chen, X. Yang, Y. Yu, M. Ma, Z. Guo, H. Yang, and B. Huang, “Object-oriented option framework for robotics manipulation in clutter,” in IROS , 2023
2023
Closest in time.
J. Pang, X. Yang, S. Yang, X. Chen, and Y. Yu, “Natural language instruction-following with task-related language development and translation,” in NeurIPS , 2023
2023
Closest in time.
X. Li, G. Chen, P. Amyotte, F. Khan, and M. Alauddin, “Vulnerability assessment of storage tanks exposed to simultaneous fire and explosion hazards,” Reliability Engineering & System Safety , vol. 230, p. 108960, 2023
2023
Closest in time.
X. Cao, C. Peng, Y. Zheng, S. Li, T. T. Ha, V. Shutyaev, V. Katsikis, and P. Stanimirovic, “Neural networks for portfolio analysis in high-frequency trading,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.
X. Cao and S. Li, “Neural networks for portfolio analysis with cardinality constraints,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.
M. Lian, Z. Guo, X. Wang, S. Wen, and T. Huang, “Adaptive exact penalty design for optimal resource allocation,” IEEE Transactions on Neural Networks and Learning Systems , vol. 34, no. 3, pp. 1430–1438, 2023
2023
Closest in time.
R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.
H. Xu, L. Jiang, J. Li, Z. Yang, Z. Wang, and X. Zhan, “Sparse q-learning: Offline reinforcement learning with implicit value regularization,” in ICLR Offline RL Workshop , 2023
2023
Closest in time.
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath, “Real-world humanoid locomotion with reinforcement learning,” Science Robotics , vol. 9, no. 89, 2024
2024
Closest in time.
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner, M. Bloesch, K. Hartikainen, A. Byravan, L. Hasenclever, Y. Tassa, F. Sadeghi, N. Batchelor, F. Casarini, S. Saliceti, C. Game, N. Sreendra, K. Patel, M. Gwira, A. Huber, N. Hurley, F. Nori, R. Hadsell, and N. Heess, “Learning agile soccer skills for a bipedal robot with deep reinforcement learning,” Science Robotics , vol. 9, no. 89, 2024
2024
Closest in time.
C. Spatharis and K. Blekas, “Multiagent reinforcement learning for autonomous driving in traffic zones with unsignalized intersections,” Journal of Intelligent Transportation Systems , vol. 28, no. 1, pp. 103–119, 2024
2024
Closest in time.
Y. Wang, J. Zhang, Y. Chen, H. Yuan, and C. Wu, “Automatic learning-based data optimization method for autonomous driving,” Digital Signal Processing , vol. 148, p. 104428, 2024
2024
Closest in time.
B. Petryshyn, S. Postupaiev, S. Ben Bari, and A. Ostreika, “Deep reinforcement learning for autonomous driving in amazon web services deepracer,” Information , vol. 15, no. 2, p. 113, 2024
2024
Closest in time.
C. Jia, F. Zhang, T. Xu, J. Pang, Z. Zhang, and Y. Yu, “Model gradient: Unified model and policy learning in model-based reinforcement learning,” Frontiers of Computer Science , vol. 18, no. 4, p. 184339, 2024
2024
Closest in time.