Fetching the paper…
Reading the bibliography…
Ideally, we would place a robot in a real-world environment and leave it there improving on its own by gathering more experience autonomously.
Understanding teacher gaze patterns for robot learning
A. Saran, E. S. Short, A. Thomaz, and S. Niekum · 1907
Earlier work this paper cites.
Deep bayesian reward learning from preferences
D. S. Brown and S. Niekum · 1912
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
S. Lange, M. Riedmiller, and A. Voigtländer · 2012
Earlier work this paper cites.
Asking for help using inverse semantics
S. Tellex, R. Knepper, A. Li, D. Rus, and N. Roy · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Learning compound multi-step controllers under unknown dynamics
W. Han, S. Levine, and P. Abbeel · 2015
Earlier work this paper cites.
Planit: A crowdsourcing approach for learning to plan paths from large scale preference feedback
A. Jain, D. Das, J. K. Gupta, and A. Saxena · 2015
Earlier work this paper cites.
Neural autoregressive distribution estimation
B. Uria, M. Côté, K. Gregor, I. Murray, and H. Larochelle · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
A. van den Oord, N. Kalchbrenner, L. Espeholt, K. Kavukcuoglu, O. Vinyals, and A. Graves · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
C-learn: Learning geometric constraints from demonstrations for multi-step manipulation in shared autonomy
C. Pérez-D’Arpino and J. A. Shah · 2017
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Learning multi-level hierarchies with hindsight
A. Levy, G. Konidaris, R. Platt, and K. Saenko · 2017
Earlier work this paper cites.
Density estimation using real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2017
Earlier work this paper cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
B. Eysenbach, S. Gu, J. Ibarz, and S. Levine · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
B. Eysenbach, S. Gu, J. Ibarz, and S. Levine · 2018
Earlier work this paper cites.
Variational inverse control with events: A general framework for data-driven reward definition
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al · 2018
Earlier work this paper cites.
Learning under misspecified objective spaces
A. Bobu, A. Bajcsy, J. F. Fisac, and A. D. Dragan · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, V. Kumar, and W. Zaremba · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Variational inverse control with events: A general framework for data-driven reward definition
Situational confidence assistance for lifelong shared autonomy
M. Zurek, A. Bobu, D. S. Brown, and A. D. Dragan · 2021
Later among the works it cites.
Wish you were here: Hindsight goal selection for long-horizon dexterous manipulation
T. Davchev, O. Sushkov, J.-B. Regli, S. Schaal, Y. Aytar, M. Wulfmeier, and J. Scholz · 2021
Later among the works it cites.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
L. M. Smith, I. Kostrikov, and S. Levine · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg · 2022
Later among the works it cites.
Reward (mis)design for autonomous driving
W. B. Knox, A. Allievi, H. Banzhaf, F. Schmitt, and P. Stone · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine · 2018
Cited alongside, same era.
Learning to reach goals without reinforcement learning
D. Ghosh, A. Gupta, J. Fu, A. Reddy, C. Devin, B. Eysenbach, and S. Levine · 2019
Cited alongside, same era.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar · 2019
Cited alongside, same era.
A data-efficient framework for training and sim-to-real transfer of navigation policies
H. Bharadhwaj, Z. Wang, Y. Bengio, and L. Paull · 2019
Cited alongside, same era.
Mo’states mo’problems: Emergency stop mechanisms from observation
S. Ainsworth, M. Barnes, and S. Srinivasa · 2019
Cited alongside, same era.
Active learning of reward dynamics from hierarchical queries
C. Basu, E. Bıyık, Z. He, M. Singhal, and D. Sadigh · 2019
Cited alongside, same era.
Dynamical distance learning for unsupervised and semi-supervised skill discovery
K. Hartikainen, X. Geng, T. Haarnoja, and S. Levine · 2019
Cited alongside, same era.
A. Handa, A. Allshire, V. Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. V. Wyk, A. Zhurkevich, B. Sundaralingam, Y. Narang, J. Lafleche, D. Fox, and G. State · 2022
Later among the works it cites.
A state-distribution matching approach to non-episodic reinforcement learning
A. Sharma, R. Ahmad, and C. Finn · 2022
Later among the works it cites.
Dexterous manipulation from images: Autonomous real-world RL via substep guidance
K. Xu, Z. Hu, R. Doshi, A. Rovinsky, V. Kumar, A. Gupta, and S. Levine · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Towards real robot learning in the wild: A case study in bipedal locomotion
M. Bloesch, J. Humplik, V. Patraucean, R. Hafner, T. Haarnoja, A. Byravan, N. Y. Siegel, S. Tunyasuvunakool, F. Casarini, N. Batchelor, et al · 2022
Later among the works it cites.
Demonstration-bootstrapped autonomous practicing via multi-task reinforcement learning
A. Gupta, C. Lynch, B. Kinman, G. Peake, S. Levine, and K. Hausman · 2022
Later among the works it cites.
Learning preferences for interactive autonomy
E. Biyik · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Physical interaction as communication: Learning robot objectives online from human corrections
D. P. Losey, A. Bajcsy, M. K. O’Malley, and A. D. Dragan · 2022
Later among the works it cites.
Revisiting human-robot teaching and learning through the lens of human concept learning
S. Booth, S. Sharma, S. Chung, J. Shah, and E. L. Glassman · 2022
Later among the works it cites.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Later among the works it cites.
Interactive language: Talking to robots in real time
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2022
Later among the works it cites.
Self-improving robots: End-to-end autonomous visuomotor reinforcement learning
A. Sharma, A. M. Ahmed, R. Ahmad, and C. Finn · 2023
Closest in time.
Self-improving robots: End-to-end autonomous visuomotor reinforcement learning
A. Sharma, A. M. Ahmed, R. Ahmad, and C. Finn · 2023
Closest in time.
When learning is out of reach, reset: Generalization in autonomous visuomotor reinforcement learning
Z. Zhang and L. Weihs · 2023
Closest in time.
Inverse preference learning: Preference-based rl without a reward function
J. Hejna and D. Sadigh · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Breadcrumbs to the goal: Goal-conditioned exploration from human-in-the-loop feedback
M. Torne, M. Balsells, Z. Wang, S. Desai, T. Chen, P. Agrawal, and A. Gupta · 2023
Closest in time.
Robots that ask for help: Uncertainty alignment for large language model planners
A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, et al · 2023
Closest in time.
Diagnosis, feedback, adaptation: A human-in-the-loop framework for test-time policy adaptation
A. Peng, A. Netanyahu, M. K. Ho, T. Shu, A. Bobu, J. Shah, and P. Agrawal · 2023
Closest in time.