Fetching the paper…
Reading the bibliography…
We propose Deep Optimistic Linear Support Learning (DOL) to solve high-dimensional multi-objective decision problems where the relative importances of the objectives are not known a priori.
Algorithms for partially observable Markov decision processes
H.-T. Cheng · 1988
Earlier work this paper cites.
Reinforcement Learning for Robots Using Neural Networks
L. Lin · 1993
Earlier work this paper cites.
Introduction to reinforcement learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Dynamic preferences in multi-criteria reinforcement learning
S. Natarajan and P. Tadepalli · 2005
Earlier work this paper cites.
Managing power consumption and performance of computing systems using reinforcement learning
G. Tesauro, R. Das, H. Chan, J. O. Kephart, C. Lefurgy, D. W. Levine, and F. Rawson · 2007
Earlier work this paper cites.
Evolutionary algorithms for solving multi-objective problems
C. C. Coello, G. B. Lamont, and D. A. Van Veldhuizen · 2007
Earlier work this paper cites.
Learning all optimal policies with multiple criteria
L. Barrett and S. Narayanan · 2008
Earlier work this paper cites.
Solving multi-objective reinforcement learning problems by EDA-RL - acquisition of various strategies
H. Handa · 2009
Earlier work this paper cites.
Empirical evaluation methods for multiobjective reinforcement learning algorithms
P. Vamplew, R. Dazeley, A. Berry, E. Dekker, and R. Issabekov · 2011
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley · 2013
Earlier work this paper cites.
Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of Pareto dominating policies
K. V. Moffaert and A. Nowé · 2014
Earlier work this paper cites.
Model-based multi-objective reinforcement learning
M. A. Wiering, M. Withagen, and M. M. Drugan · 2014
Cited alongside, same era.
The scalarized multi-objective multi-armed bandit problem: an empirical study of its exploration vs. exploitation tradeoff
S. Q. Yahyaa, M. M. Drugan, and B. Manderick · 2014
Cited alongside, same era.
A novel adaptive weight selection algorithm for multi-objective multi-agent reinforcement learning
K. Van Moffaert, T. Brys, A. Chandra, L. Esterle, P. R. Lewis, and A. Nowé · 2014
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Cited alongside, same era.
Data-efficient learning of feedback policies from image pixels using deep dynamical models
Y. M. Assael, N. Wahlström, T. B. Schön, and M. P. Deisenroth · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. D. Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, S. Legg, V. Mnih, K. Kavukcuoglu, and D. Silver · 2015
Later among the works it cites.
Move Evaluation in Go Using Deep Convolutional Neural Networks
C. J. Maddison, A. Huang, I. Sutskever, and D. Silver · 2015
Later among the works it cites.
Point-based planning for multi-objective POMDPs
D. M. Roijers, S. Whiteson, and F. A. Oliehoek · 2015
Later among the works it cites.
Pareto local policy search for MOMDP planning
C. Kooijman, M. de Waard, M. Inja, D. M. Roijers, and S. Whiteson · 2015
Later among the works it cites.
Variational multi-objective coordination
D. M. Roijers, S. Whiteson, A. T. Ihler, and F. A. Oliehoek · 2015
Later among the works it cites.
Learning to communicate with deep multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. T. Springenberg, J. Boedecker, and M. A. Riedmiller · 2015
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, and M. Lanctot · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Computing convex coverage sets for faster multi-objective coordination
D. M. Roijers, S. Whiteson, and F. A. Oliehoek
Cited in the paper.
J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson · 2016
Closest in time.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Closest in time.
Increasing the action gap: New operators for reinforcement learning
M. G. Bellemare, G. Ostrovski, A. Guez, P. S. Thomas, and R. Munos · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Closest in time.
Multi-Objective Decision-Theoretic Planning
D. M. Roijers · 2016
Closest in time.