Fetching the paper…
Reading the bibliography…
In this paper, we introduce a novel form of value function, $Q(s, s')$, that expresses the utility of transitioning from a state $s$ to a neighboring state $s'$ and then acting optimally thereafter.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S · 2003
Earlier work this paper cites.
Double q-learning
Hasselt, H. V · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Cited alongside, same era.
Intrinsically motivated goal exploration processes with automatic curriculum learning
Forestier, S., Mollard, Y., and Oudeyer, P.-Y · 2017
Cited alongside, same era.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Liu, Y., Gupta, A., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from multi-view observation
Sermanet, P., Lynch, C., Hsu, J., and Levine, S · 2017
Cited alongside, same era.
Learning what you can do before doing anything
Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A · 2018
Later among the works it cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Later among the works it cites.
Learn what not to learn: Action elimination with deep reinforcement learning
Zahavy, T., Haroush, M., Merlis, N., Mankowitz, D. J., and Mannor, S · 2018
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S · 2019
Later among the works it cites.
Learning action representations for reinforcement learning
Chandak, Y., Theocharous, G., Kostas, J., Jordan, S., and Thomas, P. S · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Forward-backward reinforcement learning
Edwards, A. D., Downs, L., and Davidson, J. C · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Recall traces: Backtracking models for efficient reinforcement learning
Goyal, A., Brakel, P., Fedus, W., Lillicrap, T., Levine, S., Larochelle, H., and Bengio, Y · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Cited alongside, same era.
Learning plannable representations with causal infogan
Kurutach, T., Tamar, A., Yang, G., Russell, S. J., and Abbeel, P · 2018
Cited alongside, same era.
Later among the works it cites.
Learning action-transferable policy with action embedding
Chen, Y., Chen, Y., Yang, Y., Li, Y., Yin, J., and Fan, C · 2019
Later among the works it cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Later among the works it cites.
Cross domain imitation learning
Kim, K. H., Gu, Y., Song, J., Zhao, S., and Ermon, S · 2019
Later among the works it cites.
State alignment-based imitation learning
Liu, F., Ling, Z., Mu, T., and Su, H · 2019
Later among the works it cites.
Mandlekar, A., Ramos, F., Boots, B., Fei-Fei, L., Garg, A., and Fox, D · 2019
Later among the works it cites.
Addressing sample complexity in visual tasks using her and hallucinatory gans
Sahni, H., Buckley, T., Abbeel, P., and Kuzovkin, I · 2019
Later among the works it cites.
Learning predictive models from observation and interaction
Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C · 2019
Later among the works it cites.
Provably efficient imitation learning from observation alone
Sun, W., Vemula, A., Boots, B., and Bagnell, J. A · 2019
Later among the works it cites.
Recent advances in imitation learning from observation
Torabi, F., Warnell, G., and Stone, P · 2019
Later among the works it cites.