Fetching the paper…
Reading the bibliography…
A fundamental trait of intelligence is the ability to achieve goals in the face of novel circumstances, such as making decisions from new action choices.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M. and Paliwal, K. K · 1997
Earlier work this paper cites.
Statistical learning theory, 1998
Vapnik, V · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Introduction to statistical learning theory
Bousquet, O., Boucheron, S., and Lugosi, G · 2003
Earlier work this paper cites.
The problem of overfitting
Hawkins, D. M · 2004
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Legg, S. and Hutter, M · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
The nature of statistical learning theory
Vapnik, V · 2013
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
Novelty and inductive generalization in human reinforcement learning
Gershman, S. J. and Niv, Y · 2015
Earlier work this paper cites.
Deep reinforcement learning in parameterized action space
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Deep reinforcement learning with a natural language action space
He, J., Chen, J., He, X., Gao, J., Li, L., Deng, L., and Ostendorf, M · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S · 2017
Cited alongside, same era.
Unsupervised learning of disentangled representations from video
Denton, E. L. and Birodkar, v · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Ebert, F., Finn, C., Lee, A. X., and Levine, S · 2017
Cited alongside, same era.
Towards a neural statistician
Edwards, H. and Storkey, A · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Later among the works it cites.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D · 2018
Later among the works it cites.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S · 2018
Later among the works it cites.
Rohde, D., Bonner, S., Dunlop, T., Vasile, F., and Karatzoglou, A · 2018
Later among the works it cites.
Graph networks as learnable physics engines for inference and control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-shot task generalization with multi-task deep reinforcement learning
Oh, J., Singh, S., Lee, H., and Kohli, P · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Robust imitation of diverse behaviors
Wang, Z., Merel, J. S., Reed, S. E., de Freitas, N., Wayne, G., and Heess, N · 2017
Cited alongside, same era.
Neural task programming: Learning to generalize across hierarchical tasks
Xu, D., Nair, S., Zhu, Y., Gao, J., Garg, A., Fei-Fei, L., and Savarese, S · 2017
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Cited alongside, same era.
Sanchez-Gonzalez, A., Heess, N., Springenberg, J. T., Merel, J., Riedmiller, M., Hadsell, R., and Battaglia, P · 2018
Later among the works it cites.
Improving generalization for abstract reasoning tasks using disentangled feature representations
Steenbrugge, X., Leroux, S., Verbelen, T., and Dhoedt, B · 2018
Later among the works it cites.
Nervenet: Learning structured policy with graph neural networks
Wang, T., Liao, R., Ba, J., and Fidler, S · 2018
Later among the works it cites.
The tools challenge: Rapid trial-and-error learning in physical problem solving
Allen, K. R., Smith, K. A., and Tenenbaum, J. B · 2019
Later among the works it cites.
Phyre: A new benchmark for physical reasoning
Bakhtin, A., van der Maaten, L., Johnson, J., Gustafson, L., and Girshick, R · 2019
Later among the works it cites.
Learning action representations for reinforcement learning
Chandak, Y., Theocharous, G., Kostas, J., Jordan, S., and Thomas, P. S · 2019
Later among the works it cites.
Learning action-transferable policy with action embedding
Chen, Y., Chen, Y., Yang, Y., Li, Y., Yin, J., and Fan, C · 2019
Later among the works it cites.
EMI: Exploration with mutual information
Kim, H., Kim, J., Jeong, Y., Levine, S., and Song, H. O · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2019
Later among the works it cites.
Learning to control self-assembling morphologies: A study of generalization via modularity
Pathak, D., Lu, C., Darrell, T., Isola, P., and Efros, A. A · 2019
Later among the works it cites.
The natural language of actions
Tennenholtz, G. and Mannor, S · 2019
Later among the works it cites.
Improvisation through physical understanding: Using novel objects as tools with visual foresight, 2019
Xie, A., Ebert, F., Levine, S., and Finn, C · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Closest in time.
Lifelong learning with a changing action set
Chandak, Y., Theocharous, G., Nota, C., and Thomas, P · 2020
Closest in time.