Fetching the paper…
Reading the bibliography…
We consider the problem of learning from sparse and underspecified rewards, where an agent receives a complex input, such as a natural language instruction, and needs to generate a complex response, such as an action sequence, while only receiving binary success-failure feedback.
Procedures as a representation for data in a computer program for understanding natural language
Winograd, T · 1971
Earlier work this paper cites.
Understanding natural language
Winograd, T · 1972
Earlier work this paper cites.
On bayesian methods for seeking the extremum
Močkus, J · 1975
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
Zelle, M. and Mooney, R. J · 1996
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Dayan, P. and Hinton, G. E · 1997
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Gaussian processes in machine learning
Rasmussen, C. E · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
Chen, D. L. and Mooney, R. J · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Artzi, Y. and Zettlemoyer, L · 2013
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Parallelizing exploration-exploitation tradeoffs in gaussian process bandit optimization
Desautels, T., Krause, A., and Burdick, J. W · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and Liang, P · 2015
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I. J., Harp, A., Irving, G., Isard, M., Jia, Y., Józefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D. G., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P. A., Vanhoucke, V., Vasudevan, V., Viégas, F. B., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
RL 2 \text{RL}^{2} : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Simpler context-dependent logical forms via model projections
Long, R., Pasupat, P., and Liang, P · 2016
Cited alongside, same era.
Learning a natural language interface with neural programmer
Neelakantan, A., Le, Q. V., Abadi, M., McCallum, A. D., and Amodei, D · 2016
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Norouzi, M., Bengio, S., Jaitly, N., Schuster, M., Wu, Y., Schuurmans, D., et al · 2016
Cited alongside, same era.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Cited alongside, same era.
The value of semantic parse labeling for knowledge base question answering
Yih, W.-t., Richardson, M., Meek, C., Chang, M.-W., and Suh, J · 2016
Cited alongside, same era.
Multi-task maximum entropy inverse reinforcement learning
Gleave, A. and Habryka, O · 2018
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Neural multi-step reasoning for question answering on semi-structured tables
Haug, T., Ganea, O.-E., and Grnarova, P · 2018
Later among the works it cites.
Natural language to structured query generation via meta-learning
Huang, P.-S., Wang, C., Singh, R., tau Yih, W., and He, X · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reward shaping via meta-learning
Zou, H., Ren, T., Yan, D., Su, H., and Zhu, J · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Google vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D · 2017
Cited alongside, same era.
From language to programs: Bridging reinforcement learning and maximum marginal likelihood
Guu, K., Pasupat, P., Liu, E., and Liang, P · 2017
Cited alongside, same era.
Grounded language learning in a simulated 3d world
Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., et al · 2017
Cited alongside, same era.
Neural semantic parsing with type constraints for semi-structured tables
Krishnamurthy, J., Dasigi, P., and Gardner, M · 2017
Cited alongside, same era.
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Memory augmented policy optimization for program synthesis and semantic parsing
Liang, C., Norouzi, M., Berant, J., Le, Q. V., and Lao, N · 2018
Later among the works it cites.
It was the training data pruning too!
Mudrakarta, P. K., Taly, A., Sundararajan, M., and Dhamdhere, K · 2018
Later among the works it cites.
Deep online learning via meta-learning: Continual adaptation for model-based rl
Nagabandi, A., Finn, C., and Levine, S · 2018
Later among the works it cites.
Reptile: a scalable metalearning algorithm
Nichol, A. and Schulman, J · 2018
Later among the works it cites.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R · 2018
Later among the works it cites.
Towards diverse text generation with inverse reinforcement learning
Shi, Z., Chen, X., Qiu, X., and Huang, X · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
No metrics are perfect: Adversarial reward learning for visual storytelling
Wang, X., Chen, W., Wang, Y.-F., and Wang, W. Y · 2018
Later among the works it cites.
Learning to teach with dynamic loss functions
Wu, L., Tian, F., Xia, Y., Fan, Y., Qin, T., Jian-Huang, L., and Liu, T.-Y · 2018
Later among the works it cites.
Few-shot goal inference for visuomotor learning and planning
Xie, A., Singh, A., Levine, S., and Finn, C · 2018
Later among the works it cites.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Later among the works it cites.
Metric-optimized example weights
Zhao, S., Fard, M. M., and Gupta, M · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Zheng, Z., Oh, J., and Singh, S · 2018
Later among the works it cites.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E · 2019
Closest in time.
Explaining relational queries to non-experts
Berant, J., Deutch, D., Globerson, A., Milo, T., and Wolfson, T · 2019
Closest in time.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Fu, J., Korattikara, A., Levine, S., and Guadarrama, S · 2019
Closest in time.
Self-supervised generalisation with meta auxiliary learning
Liu, S., Davison, A. J., and Johns, E · 2019
Closest in time.