Fetching the paper…
Reading the bibliography…
Many reinforcement learning tasks can benefit from explicit planning based on an internal model of the environment.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Emergence of scaling in random networks
Barabási, A.-L. and Albert, R · 1999
Earlier work this paper cites.
Networks, dynamics, and the small-world phenomenon
Watts, D. J · 1999
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Yanardag, P. and Vishwanathan, S · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Neural message passing for quantum chemistry
Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E · 2017
Cited alongside, same era.
Learning model-based planning from scratch
Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racanière, S., Reichert, D., Weber, T., Wierstra, D., and Battaglia, P · 2017
Cited alongside, same era.
Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y · 2017
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Later among the works it cites.
Neural execution of graph algorithms
Veličković, P., Ying, R., Padovano, M., Hadsell, R., and Blundell, C · 2019
Later among the works it cites.
What can neural networks reason about?
Xu, K., Li, J., Zhang, M., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S · 2019
Later among the works it cites.
Principal neighbourhood aggregation for graph nets
Corso, G., Cavalleri, L., Beaini, D., Liò, P., and Veličković, P · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Niu, S., Chen, S., Guo, H., Targonski, C., Smith, M. C., and Kovačević, J · 2018
Cited alongside, same era.
On random graphs
Erdös, P. et al
Cited in the paper.
Rivlin, O., Hazan, T., and Karpas, E · 2020
Closest in time.
Neural Execution Engines, 2020
Yan, Y., Swersky, K., Koutra, D., Ranganathan, P., and Hashemi, M · 2020
Closest in time.