Fetching the paper…
Reading the bibliography…
Many real-world problems require trading off multiple competing objectives.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Multi-criteria reinforcement learning
Gábor, Z., Kalmár, Z., and Szepesvári, C · 1998
Earlier work this paper cites.
Q-decomposition for reinforcement learning agents
Russell, S. and Zimdars, A. L · 2003
Earlier work this paper cites.
Function-transformation methods for multi-objective optimization
Marler, R. T. and Arora, J. S · 2005
Earlier work this paper cites.
Dynamic preferences in multi-criteria reinforcement learning
Natarajan, S. and Tadepalli, P · 2005
Earlier work this paper cites.
Normalization and other topics in multi-objective optimization
Grodzevich, O. and Romanko, O · 2006
Earlier work this paper cites.
Adaptive weighted sum method for multiobjective optimization: A new method for pareto front generation
Kim, I. and de Weck, O · 2006
Earlier work this paper cites.
Learning all optimal policies with multiple criteria
Barrett, L. and Narayanan, S · 2008
Earlier work this paper cites.
Empirical evaluation methods for multiobjective reinforcement learning algorithms
Vamplew, P., Dazeley, R., Berry, A., Issabekov, R., and Dekker, E · 2011
Earlier work this paper cites.
DEAP: Evolutionary algorithms made easy
Fortin, F.-A., De Rainville, F.-M., Gardner, M.-A., Parizeau, M., and Gagné, C · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R · 2013
Earlier work this paper cites.
Scalarized multi-objective reinforcement learning: Novel design techniques
Van Moffaert, K., Drugan, M. M., and Nowé, A · 2013
Earlier work this paper cites.
New perspectives on the natural gradient method
Martens, J · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of Pareto dominating policies
Moffaert, K. V. and Nowé, A · 2014
Earlier work this paper cites.
Policy gradient approaches for multi-objective sequential decision making
Parisi, S., Pirotta, M., Smacchia, N., Bascetta, L., and Restelli, M · 2014
Earlier work this paper cites.
Linear support for multi-objective coordination graphs
Roijers, D. M., Whiteson, S., and Oliehoek, F. A · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Multiobjective reinforcement learning: A comprehensive overview
Liu, C., Xu, X., and Hu, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Multi-objective reinforcement learning with continuous Pareto frontier approximation
Pirotta, M., Parisi, S., and Restelli, M · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Multi-objective MDPs with conditional lexicographic reward preferences
Wray, K. H., Zilberstein, S., and Mouaddib, A.-I · 2015
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Later among the works it cites.
Learning by playing - solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.
Sim-to-real: Learning agile locomotion for quadruped robots
Tan, J., Zhang, T., Coumans, E., Iscen, A., Bai, Y., Hafner, D., Bohez, S., and Vanhoucke, V · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M · 2018
Later among the works it cites.
Dynamic weights in multi-objective deep reinforcement learning
Abels, A., Roijers, D. M., Lenaerts, T., Nowé, A., and Steckelmacher, D · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mossalam, H., Assael, Y. M., Roijers, D. M., and Whiteson, S · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Cited alongside, same era.
Learning values across many orders of magnitude
van Hasselt, H. P., Guez, A., Hessel, M., Mnih, V., and Silver, D · 2016
Cited alongside, same era.
ϵ \epsilon -pal: An active learning approach to the multi-objective optimization problem
Zuluaga, M., Krause, A., and Püschel, M · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
A robust proportion-preserving composite objective function for scale-invariant multi-objective optimization
Daneshmand, M., Tale Masouleh, M., Saadatzi, M. H., Ozcinar, C., and Anbarjafari, G · 2017
Cited alongside, same era.
Later among the works it cites.
Value constrained model-free continuous control
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R · 2019
Later among the works it cites.
Meta-learning for multi-objective reinforcement learning
Chen, X., Ghadirzadeh, A., Björkman, M., and Jensfelt, P · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2019
Later among the works it cites.
Learning gentle object manipulation with curiosity-driven deep reinforcement learning
Huang, S. H., Zambelli, M., Kay, J., Martins, M. F., Tassa, Y., Pilarski, P. M., and Hadsell, R · 2019
Later among the works it cites.
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y · 2019
Later among the works it cites.
Hierarchical visuomotor control of humanoids
Merel, J., Ahuja, A., Pham, V., Tunyasuvunakool, S., Liu, S., Tirumala, D., Heess, N., and Wayne, G · 2019
Later among the works it cites.
Using logical specifications of objectives in multi-objective reinforcement learning
Nottingham, K., Balakrishnan, A., Deshmukh, J., Christopherson, C., and Wingate, D · 2019
Later among the works it cites.
Pareto-DQN: Approximating the Pareto front in complex multi-objective decision problems
Reymond, M. and Nowé, A · 2019
Later among the works it cites.
Regularized hierarchical policies for compositional transfer in robotics
Wulfmeier, M., Abdolmaleki, A., Hafner, R., Springenberg, J. T., Neunert, M., Hertweck, T., Lampe, T., Siegel, N., Heess, N., and Riedmiller, M · 2019
Later among the works it cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Yang, R., Sun, X., and Narasimhan, K · 2019
Later among the works it cites.
Tossingbot: Learning to throw arbitrary objects with residual physics
Zeng, A., Song, S., Lee, J., Rodriguez, A., and Funkhouser, T · 2019
Later among the works it cites.
Zhan, H. and Cao, Y · 2019
Later among the works it cites.
CVXOPT: A Python package for convex optimization, version 1.2
Andersen, M. S., Dahl, J., and Vandenberghe, L · 2020
Closest in time.
CoMic: Co-Training and Mimicry for Reusable Skills
Anonymous · 2020
Closest in time.
Shadow dexterous hand
Shadow Robot Company · 2020
Closest in time.
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., Heess, N., Belov, D., Riedmiller, M., and Botvinick, M. M · 2020
Closest in time.