Fetching the paper…
Reading the bibliography…
Multi-objective reinforcement learning (MORL) is the generalization of standard reinforcement learning (RL) approaches to solve sequential decision making problems that consist of several, possibly conflicting, objectives.
I. Das and J. E. Dennis, “A closer look at drawbacks of minimizing weighted sums of objectives for pareto set generation in multicriteria optimization problems,” Structural optimization , vol. 14, no. 1, pp. 63–69, 1997
1997
Earlier work this paper cites.
K. Deb, S. Agrawal, A. Pratap, and T. Meyarivan, “A fast elitist non-dominated sorting genetic algorithm for multi-objective optimization: Nsga-ii,” in International Conference on Parallel Problem Solving From Nature . Springer, 2000, pp. 849–858
2000
Earlier work this paper cites.
K. Miettinen, “Some methods for nonlinear multi-objective optimization,” in International Conference on Evolutionary Multi-Criterion Optimization . Springer, 2001, pp. 1–20
2001
Earlier work this paper cites.
E. Zitzler, M. Laumanns, and L. Thiele, “Spea2: Improving the strength pareto evolutionary algorithm,” TIK-report , vol. 103, 2001
2001
Earlier work this paper cites.
K. Miettinen and M. M. Mäkelä, “On scalarizing functions in multiobjective optimization,” OR spectrum , vol. 24, no. 2, pp. 193–213, 2002
2002
Earlier work this paper cites.
S. Natarajan and P. Tadepalli, “Dynamic preferences in multi-criteria reinforcement learning,” in Proceedings of the 22nd international conference on Machine learning . ACM, 2005, pp. 601–608
2005
Earlier work this paper cites.
A. Konak, D. W. Coit, and A. E. Smith, “Multi-objective optimization using genetic algorithms: A tutorial,” Reliability Engineering & System Safety , vol. 91, no. 9, pp. 992–1007, 2006
2006
Earlier work this paper cites.
C. A. C. Coello, G. B. Lamont, D. A. Van Veldhuizen et al. , Evolutionary algorithms for solving multi-objective problems . Springer, 2007, vol. 5
2007
Earlier work this paper cites.
H. Handa, “Eda-rl: estimation of distribution algorithms for reinforcement learning problems,” in Proceedings of the 11th Annual conference on Genetic and evolutionary computation . ACM, 2009, pp. 405–412
2009
Earlier work this paper cites.
——, “Solving multi-objective reinforcement learning problems by eda-rl-acquisition of various strategies,” in 2009 ninth international conference on intelligent systems design and applications . IEEE, 2009, pp. 426–431
2009
Earlier work this paper cites.
D. J. Lizotte, M. H. Bowling, and S. A. Murphy, “Efficient reinforcement learning with multiple reward functions for randomized controlled trial analysis,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10) . Citeseer, 2010, pp. 695–702
2010
Cited alongside, same era.
P. Vamplew, R. Dazeley, A. Berry, R. Issabekov, and E. Dekker, “Empirical evaluation methods for multiobjective reinforcement learning algorithms,” Machine learning , vol. 84, no. 1-2, pp. 51–80, 2011
2011
Cited alongside, same era.
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley, “A survey of multi-objective sequential decision-making,” Journal of Artificial Intelligence Research , vol. 48, pp. 67–113, 2013
2013
Cited alongside, same era.
K. Van Moffaert, M. M. Drugan, and A. Nowé, “Scalarized multi-objective reinforcement learning: Novel design techniques.” in ADPRL , 2013, pp. 191–199
2013
Cited alongside, same era.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in International Conference on Machine Learning , 2016, pp. 1329–1338
2016
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in Advances in Neural Information Processing Systems , 2017, pp. 5048–5058
2017
Later among the works it cites.
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman, “Deep predictive policy training using reinforcement learning,” in Intelligent Robots and Systems (IROS), 2017 IEEE/RSJ International Conference on . IEEE, 2017, pp. 2351–2358
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Van Moffaert, M. M. Drugan, and A. Nowé, “Hypervolume-based multi-objective reinforcement learning,” in International Conference on Evolutionary Multi-Criterion Optimization . Springer, 2013, pp. 352–366
2013
Cited alongside, same era.
S. Parisi, M. Pirotta, N. Smacchia, L. Bascetta, and M. Restelli, “Policy gradient approaches for multi-objective sequential decision making,” in Neural networks (ijcnn), 2014 international joint conference on . IEEE, 2014, pp. 2323–2330
2014
Cited alongside, same era.
C. Liu, X. Xu, and D. Hu, “Multiobjective reinforcement learning: A comprehensive overview,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , vol. 45, no. 3, pp. 385–398, 2015
2015
Cited alongside, same era.
M. Pirotta, S. Parisi, and M. Restelli, “Multi-objective reinforcement learning with continuous pareto frontier approximation,” in 29th AAAI Conference on Artificial Intelligence, AAAI 2015 and the 27th Innovative Applications of Artificial Intelligence Conference, IAAI 2015 . AAAI Press, 2015, pp. 2928–2934
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning , 2015, pp. 1889–1897
2015
Cited alongside, same era.
“Roboschool: open-source software for robot simulation,” https://blog.openai.com/roboschool/
Cited in the paper.
2017
Later among the works it cites.
T. Brys, A. Harutyunyan, P. Vrancx, A. Nowé, and M. E. Taylor, “Multi-objectivization and ensembles of shapings in reinforcement learning,” Neurocomputing , vol. 263, pp. 48–59, 2017
2017
Later among the works it cites.
S. Parisi, M. Pirotta, and J. Peters, “Manifold-based multi-objective policy search with sample reuse,” Neurocomputing , vol. 263, pp. 3–14, 2017
2017
Later among the works it cites.
2018
Closest in time.
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, V. Kumar, and W. Zaremba, “Multi-goal reinforcement learning: Challenging robotics environments and request for research,” 2018
2018
Closest in time.