Fetching the paper…
Reading the bibliography…
Conveying complex objectives to reinforcement learning (RL) agents can often be difficult, involving meticulous design of reward functions that are sufficiently informative yet easy enough to provide.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
The foundations of statistics
Savage, L. J · 1972
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Beirlant, J., Dudewicz, E. J., Györfi, L., and Van der Meulen, E. C · 1997
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H., Misra, N., Hnizdo, V., Fedorowicz, A., and Demchuk, E · 2003
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
Kohl, N. and Stone, P · 2004
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Learning collaborative manipulation tasks by demonstration using a haptic interface
Calinon, S., Evrard, P., Gribovskaya, E., Billard, A., and Kheddar, A · 2009
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W. B. and Stone, P · 2009
Earlier work this paper cites.
Robot motor skill coordination with EM-based reinforcement learning
Kormushev, P., Calinon, S., and Caldwell, D · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, J. and Peters, J · 2011
Earlier work this paper cites.
Online movement adaptation based on previous sensor experiences
Pastor, P., Righetti, L., Kalakrishnan, M., and Schaal, S · 2011
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
Pilarski, P. M., Dawson, M. R., Degris, T., Fahimi, F., Carey, J. P., and Sutton, R. S · 2011
Earlier work this paper cites.
Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective
Akgun, B., Cakmak, M., Yoo, J., and Thomaz, A · 2012
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Akrour, R., Schoenauer, M., and Sebag, M · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Sugiyama, H., Meguro, T., and Minami, Y · 2012
Earlier work this paper cites.
Monte carlo methods for preference learning
Viappiani, P · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Wilson, A., Fern, A., and Tadepalli, P · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
Wirth, C. and Fürnkranz, J · 2013
Earlier work this paper cites.
Active reward learning
Daniel, C., Viering, M., Metz, J., Kroemer, O., and Peters, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Hierarchical relative entropy policy search
Daniel, C., Neumann, G., Kroemer, O., and Peters, J · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Levine, S., Pastor, P., Krizhevsky, A., Ibarz, J., and Quillen, D · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Later among the works it cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Warnell, G., Waytowich, N., Lawhern, V., and Stone, P · 2018
Later among the works it cites.
Few-shot goal inference for visuomotor learning and planning
Xie, A., Singh, A., Levine, S., and Finn, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Pinto, L. and Gupta, A · 2016
Cited alongside, same era.
Model-free preference-based reinforcement learning
Wirth, C., Fürnkranz, J., and Neumann, G · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Roberts, D., Taylor, M. E., and Littman, M. L · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Cited alongside, same era.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Zhang, T., McCarthy, Z., Jow, O., Lee, D., Goldberg, K., and Abbeel, P · 2018
Later among the works it cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Later among the works it cites.
Deep reinforcement learning from policy-dependent human feedback
Arumugam, D., Lee, J. K., Saskin, S., and Littman, M. L · 2019
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Later among the works it cites.
Preferences implicit in the state of the world
Shah, R., Krasheninnikov, D., Alexander, J., Abbeel, P., and Dragan, A · 2019
Later among the works it cites.
End-to-end robotic reinforcement learning without reward engineering
Singh, A., Yang, L., Hartikainen, K., Finn, C., and Levine, S · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Solar: Deep structured representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Later among the works it cites.
Active preference-based gaussian process regression for reward learning
Biyik, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D · 2020
Later among the works it cites.
Human preference scaling with demonstrations for deep reinforcement learning
Cao, Z., Wong, K., and Lin, C.-T · 2020
Later among the works it cites.
Learning agile robotic locomotion skills by imitating animals
Peng, X. B., Coumans, E., Zhang, T., Lee, T.-W., Tan, J., and Levine, S · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2020
Later among the works it cites.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
Smith, L., Dhawan, N., Zhang, M., Abbeel, P., and Levine, S · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N · 2020
Later among the works it cites.
Avoiding side effects in complex environments
Turner, A. M., Ratzlaff, N., and Tadepalli, P · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
Hao, L. and Pieter, A · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Closest in time.