Fetching the paper…
Reading the bibliography…
We study reinforcement learning (RL) with no-reward demonstrations, a setting in which an RL agent has access to additional data from the interaction of other agents with the same environment.
Pattern-recognizing control systems, 1964
Widrow, B. and Smith, F. W · 1964
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1989
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Evolution of smart-n players
Stahl, D. O · 1993
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Is it an agent, or just a program?: A taxonomy for autonomous agents
Franklin, S. and Graesser, A · 1996
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
Solving a huge number of simular tasks: a combination of multi-task learning and a hierarchical bayesian approach
Heskes, T · 1998
Earlier work this paper cites.
Using artificial neural networks to model opponents in texas hold’em
Davidson, A · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
Maximum margin planning
Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Evolving explicit opponent models in game playing
Lockett, A. J., Chen, C. L., and Miikkulainen, R · 2007
Earlier work this paper cites.
Survey: Robot programming by demonstration
Billard, A., Calinon, S., Dillmann, R., and Schaal, S · 2008
Earlier work this paper cites.
Learning for control from multiple demonstrations
Coates, A., Abbeel, P., and Ng, A. Y · 2008
Earlier work this paper cites.
Game theory of mind
Yoshida, W., Dolan, R. J., and Friston, K. J · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Enhanced intelligent driver model to access the impact of driving strategies on traffic capacity
Kesting, A., Treiber, M., and Helbing, D · 2010
Earlier work this paper cites.
Inverse reinforcement learning in partially observable environments
Choi, J. and Kim, K.-E · 2011
Earlier work this paper cites.
Bayesian multitask inverse reinforcement learning
Dimitrakakis, C. and Rothkopf, C. A · 2011
Earlier work this paper cites.
Donut as i do: Learning from failed demonstrations
Grollman, D. H. and Billard, A · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
Taylor, M. E., Suay, H. B., and Chernova, S · 2011
Earlier work this paper cites.
A cascaded supervised learning approach to inverse reinforcement learning
Klein, E., Piot, B., Geist, M., and Pietquin, O · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Learning to select and generalize striking movements in robot table tennis
Mülling, K., Kober, J., Kroemer, O., and Peters, J · 2013
Earlier work this paper cites.
Learning compact parameterized skills with a single regression
Stulp, F., Raiola, G., Hoarau, A., Ivaldi, S., and Sigaud, O · 2013
Earlier work this paper cites.
Multi-task policy search for robotics
Deisenroth, M. P., Englert, P., Peters, J., and Fox, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Robust bayesian inverse reinforcement learning with sparse behavior noise
Zheng, J., Liu, S., and Ni, L. M · 2014
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., Ondruska, P., and Posner, I · 2015
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He, H., Boyd-Graber, J., Kwok, K., and Daumé III, H · 2016
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Later among the works it cites.
One-shot high-fidelity imitation: Training large-scale deep nets with rl
Paine, T. L., Colmenarejo, S. G., Wang, Z., Reed, S., Aytar, Y., Pfaff, T., Hoffman, M. W., Barth-Maron, G., Cabi, S., Budden, D., et al · 2018
Later among the works it cites.
Rabinowitz, N. C., Perbet, F., Song, H. F., Zhang, C., Eslami, S., and Botvinick, M · 2018
Later among the works it cites.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
Rahmatizadeh, R., Abolghasemi, P., Bölöni, L., and Levine, S · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Cited alongside, same era.
Inverse reinforcement learning from failure
Shiarlis, K., Messias, J., and Whiteson, S · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Observational learning by reinforcement learning
Borsa, D., Piot, B., Munos, R., and Pietquin, O · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Cited alongside, same era.
The secret of our success: How culture is driving human evolution, domesticating our species, and making us smarter
Henrich, J · 2017
Cited alongside, same era.
Later among the works it cites.
Multiple interactions made easy (mime): Large scale demonstrations data for imitation
Sharma, P., Mohan, L., Pinto, L., and Gupta, A · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Later among the works it cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Zhang, T., McCarthy, Z., Jow, O., Lee, D., Chen, X., Goldberg, K., and Abbeel, P · 2018
Later among the works it cites.
The option keyboard: Combining skills in reinforcement learning
Barreto, A., Borsa, D., Hou, S., Comanici, G., Aygün, E., Hamel, P., Toyama, D., Mourad, S., Silver, D., Precup, D., et al · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S · 2019
Later among the works it cites.
BabyAI: First steps towards grounded language learning with a human in the loop
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y · 2019
Later among the works it cites.
Robust learning from demonstrations with mixed qualities using leveraged gaussian processes
Choi, S., Lee, K., and Oh, S · 2019
Later among the works it cites.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Fu, J., Korattikara, A., Levine, S., and Guadarrama, S · 2019
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2019
Later among the works it cites.
Agent modeling as auxiliary task for deep reinforcement learning
Hernandez-Leal, P., Kartal, B., and Taylor, M. E · 2019
Later among the works it cites.
Successor uncertainties: exploration and uncertainty in temporal difference learning
Janz, D., Hron, J., Mazur, P., Hofmann, K., Hernández-Lobato, J. M., and Tschiatschek, S · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N · 2019
Later among the works it cites.
Truly batch apprenticeship learning with deep successor features
Lee, D., Srinivasan, S., and Doshi-Velez, F · 2019
Later among the works it cites.
Making efficient use of demonstrations to solve hard exploration problems
Paine, T. L., Gulcehre, C., Shahriari, B., Denil, M., Hoffman, M., Soyer, H., Tanburn, R., Kapturowski, S., Rabinowitz, N., Williams, D., et al · 2019
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Reddy, S., Dragan, A. D., and Levine, S · 2019
Later among the works it cites.
The DeepMind JAX Ecosystem, 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hennigan, T., Hessel, M., Kapturowski, S., Keck, T., Kemaev, I., King, M., Martens, L., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stokowiec, W., and Viola, F · 2020
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
Barreto, A., Hou, S., Borsa, D., Silver, D., and Precup, D · 2020
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Later among the works it cites.
Can autonomous vehicles identify, recover from, and adapt to distribution shifts?
Filos, A., Tigkas, P., McAllister, R., Rhinehart, N., Levine, S., and Gal, Y · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Hennigan, T., Cai, T., Norman, T., and Babuschkin, I · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Hoffman, M., Shahriari, B., Aslanides, J., Barth-Maron, G., Behbahani, F., Norman, T., Abdolmaleki, A., Cassirer, A., Yang, F., Baumli, K., et al · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Machado, M. C., Bellemare, M. G., and Bowling, M · 2020
Later among the works it cites.
Multi-agent social reinforcement learning improves generalization
Ndousse, K., Eck, D., Levine, S., and Jaques, N · 2020
Later among the works it cites.
Deep imitative models for flexible inference, planning, and control
Rhinehart, N., McAllister, R., and Levine, S · 2020
Later among the works it cites.