Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) enables agents to take decision based on a reward function.
D. E. Goldberg and J. H. Holland, “Genetic algorithms and machine learning,” Machine learning , vol. 3, no. 2, pp. 95–99, 1988
1988
Earlier work this paper cites.
G. Syswerda, “Uniform crossover in genetic algorithms,” in Proceedings of the third international conference on Genetic algorithms . Morgan Kaufmann Publishers, 1989, pp. 2–9
1989
Earlier work this paper cites.
L. Davis, “Handbook of genetic algorithms,” 1991
1991
Earlier work this paper cites.
D. E. Goldberg and K. Deb, “A comparative analysis of selection schemes used in genetic algorithms,” in Foundations of genetic algorithms . Elsevier, 1991, vol. 1, pp. 69–93
1991
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
J. H. Holland, “Genetic algorithms,” Scientific american , vol. 267, no. 1, pp. 66–73, 1992
1992
Earlier work this paper cites.
L. C. Baird, “Reinforcement learning in continuous time: Advantage updating,” in Neural Networks, 1994. IEEE World Congress on Computational Intelligence., 1994 IEEE International Conference on , vol. 4. IEEE, 1994, pp. 2448–2453
1994
Earlier work this paper cites.
S. Mikami and Y. Kakazu, “Genetic reinforcement learning for cooperative traffic signal control,” in Evolutionary Computation, 1994. IEEE World Congress on Computational Intelligence., Proceedings of the First IEEE Conference on . IEEE, 1994, pp. 223–228
1994
Earlier work this paper cites.
P. W. Poon and J. N. Carter, “Genetic algorithm crossover operators for ordering applications,” Computers & Operations Research , vol. 22, no. 1, pp. 135–147, 1995
1995
Earlier work this paper cites.
C. Gaskett, D. Wettergreen, and A. Zelinsky, “Q-learning in continuous state and action spaces,” in Australasian Joint Conference on Artificial Intelligence . Springer, 1999, pp. 417–428
1999
Earlier work this paper cites.
D. E. Moriarty, A. C. Schultz, and J. J. Grefenstette, “Evolutionary algorithms for reinforcement learning,” Journal of Artificial Intelligence Research , vol. 11, pp. 241–276, 1999
1999
Earlier work this paper cites.
K. Doya, “Reinforcement learning in continuous time and space,” Neural computation , vol. 12, no. 1, pp. 219–245, 2000
2000
Earlier work this paper cites.
K. Deb, A. Pratap, S. Agarwal, and T. Meyarivan, “A fast and elitist multiobjective genetic algorithm: Nsga-ii,” IEEE transactions on evolutionary computation , vol. 6, no. 2, pp. 182–197, 2002
2002
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in Robotics and Automation, 2004. Proceedings. ICRA’04. 2004 IEEE International Conference on , vol. 3. IEEE, 2004, pp. 2619–2624
2004
Cited alongside, same era.
H. V. Hasselt and M. A. Wiering, “Reinforcement learning in continuous action spaces,” 2007
2007
Cited alongside, same era.
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning cpg-based biped locomotion with a policy gradient method: Application to a humanoid robot,” The International Journal of Robotics Research , vol. 27, no. 2, pp. 213–228, 2008
2008
Cited alongside, same era.
C.-K. Lin, “H∞ reinforcement learning control of robot manipulators using fuzzy wavelet networks,” Fuzzy Sets and Systems , vol. 160, no. 12, pp. 1765–1786, 2009
2009
Cited alongside, same era.
2015
Later among the works it cites.
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in International Conference on Machine Learning , 2016, pp. 2829–2838
2016
Later among the works it cites.
A. D. Dang, H. M. La, and J. Horn, “Distributed formation control for autonomous robots following desired shapes in noisy environment,” in 2016 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI) , Sep. 2016, pp. 285–290
2016
Later among the works it cites.
Q. Wei, F. L. Lewis, Q. Sun, P. Yan, and R. Song, “Discrete-time deterministic q q -learning: A novel convergence analysis,” IEEE transactions on cybernetics , vol. 47, no. 5, pp. 1224–1237, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Liu and G. Zeng, “Study of genetic algorithm with reinforcement learning to solve the tsp,” Expert Systems with Applications , vol. 36, no. 3, pp. 6995–7001, 2009
2009
Cited alongside, same era.
J. Peters, K. Mülling, and Y. Altun, “Relative entropy policy search.” in AAAI . Atlanta, 2010, pp. 1607–1612
2010
Cited alongside, same era.
M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal, “Learning force control policies for compliant manipulation,” in Intelligent Robots and Systems (IROS), 2011 IEEE/RSJ International Conference on . IEEE, 2011, pp. 4639–4644
2011
Cited alongside, same era.
M. P. Deisenroth, C. E. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” 2011
2011
Cited alongside, same era.
M. Duguleana, F. G. Barbuceanu, A. Teirelbar, and G. Mogan, “Obstacle avoidance of redundant manipulators using neural networks based reinforcement learning,” Robotics and Computer-Integrated Manufacturing , vol. 28, no. 2, pp. 132–146, 2012
2012
Cited alongside, same era.
Z. Miljković, M. Mitić, M. Lazarević, and B. Babić, “Neural network reinforcement learning for visual control of robot manipulators,” Expert Systems with Applications , vol. 40, no. 5, pp. 1721–1736, 2013
2013
Cited alongside, same era.
H. M. La, R. S. Lim, W. Sheng, and J. Chen, “Cooperative flocking and learning in multi-robot systems for predator avoidance,” in 2013 IEEE International Conference on Cyber Technology in Automation, Control and Intelligent Systems , May 2013, pp. 337–342
2013
Cited alongside, same era.
H. M. La, R. Lim, and W. Sheng, “Multirobot cooperative learning for predator avoidance,” IEEE Transactions on Control Systems Technology , vol. 23, no. 1, pp. 52–63, Jan 2015
2015
Cited alongside, same era.
L. Jin, S. Li, H. M. La, and X. Luo, “Manipulability optimization of redundant manipulators using dynamic neural networks,” IEEE Transactions on Industrial Electronics , vol. 64, no. 6, pp. 4710–4720, June 2017
2017
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in Advances in Neural Information Processing Systems , 2017, pp. 5048–5058
2017
Later among the works it cites.
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov, “Openai baselines,” https://github.com/openai/baselines, 2017
2017
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
——, “Reinforcement learning for autonomous uav navigation using function approximation,” in 2018 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR) , Aug 2018, pp. 1–6
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Rahimi, S. Gibb, Y. Shen, and H. M. La, “A comparison of various approaches to reinforcement learning algorithms for multi-robot box pushing,” in International Conference on Engineering Research and Applications . Springer, 2018, pp. 16–30
2018
Later among the works it cites.
H. Nguyen and H. M. La, “Review of deep reinforcement learning for robot manipulation,” in The Third IEEE International Conference on Robotic Computing (IRC2019) , 2019, pp. 1–6
2019
Closest in time.