Fetching the paper…
Reading the bibliography…
Many tasks in control, robotics, and planning can be specified using desired goal configurations for various entities in the environment.
Feudal reinforcement learning. nips’93 (pp. 271–278), 1993
Dayan, P. and Hinton, G · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Neuro-dynamic Programming
Bertsekas, D. and Tsitsiklis, J · 1996
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R., and Smola, A · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2018
Cited alongside, same era.
Rearrangement: A challenge for embodied ai
Batra, D., Chang, A. X., Chernova, S., Davison, A. J., Deng, J., Koltun, V., Levine, S., Malik, J., Mordatch, I., Mottaghi, R., et al · 2020
Later among the works it cites.
Deep sets for generalization in rl
Karch, T., Colas, C., Teodorescu, L., Moulin-Frier, C., and Oudeyer, P.-Y · 2020
Later among the works it cites.
Towards practical multi-object manipulation using relational reinforcement learning
Li, R., Jabri, A., Darrell, T., and Agrawal, P · 2020
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Learning object-centric representations of multi-object scenes from multiple views
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Relational deep reinforcement learning
Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., et al · 2018
Cited alongside, same era.
Structured agents for physical construction
Bapst, V., Sanchez-Gonzalez, A., Doersch, C., Stachenfeld, K., Kohli, P., Battaglia, P., and Hamrick, J · 2019
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B · 2019
Cited alongside, same era.
Neural task graphs: Generalizing to unseen tasks from a single video demonstration
Huang, D.-A., Nair, S., Xu, D., Zhu, Y., Garg, A., Fei-Fei, L., Savarese, S., and Niebles, J. C · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Jiang, Y., Gu, S., Murphy, K., and Finn, C · 2019
Cited alongside, same era.
Contrastive learning of structured world models
Kipf, T., van der Pol, E., and Welling, M · 2019
Cited alongside, same era.
Nanbo, L., Eastwood, C., and Fisher, R. B · 2020
Later among the works it cites.
D2rl: Deep dense architectures in reinforcement learning
Sinha, S., Bharadhwaj, H., Srinivas, A., and Garg, A · 2020
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
Veerapaneni, R., Co-Reyes, J. D., Chang, M., Janner, M., Finn, C., Wu, J., Tenenbaum, J., and Levine, S · 2020
Later among the works it cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Bronstein, M. M., Bruna, J., Cohen, T., and Veličković, P · 2021
Later among the works it cites.
Roma: A relational, object-model learning agent for sample-efficient reinforcement learning
Carvalho, W., Liang, A., Lee, K., Sohn, S., Lee, H., Lewis, R. L., and Singh, S · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
The sensory neuron as a transformer: Permutation-invariant neural networks for reinforcement learning
Tang, Y. and Ha, D · 2021
Later among the works it cites.
Efficient and interpretable robot manipulation with graph neural networks
Lin, Y., Wang, A. S., Undersander, E., and Rai, A · 2022
Closest in time.