Fetching the paper…
Reading the bibliography…
Learning reward functions remains the bottleneck to equip a robot with a broad repertoire of skills.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical Analysis
Luce, R · 1959
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Stable dynamic walking over uneven terrain
Manchester, I. R., Mettin, U., Iida, F., and Tedrake, R · 2011
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Earlier work this paper cites.
Learning reward functions by integrating human demonstrations and preferences
Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D · 2019
Earlier work this paper cites.
Quadrupedal locomotion on uneven terrain with sensorized feet
Valsecchi, G., Grandia, R., and Hutter, M · 2020
Earlier work this paper cites.
Imitation learning as f-divergence minimization
Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S · 2021
Cited alongside, same era.
Lee, K., Smith, L., and Abbeel, P · 2021
Cited alongside, same era.
rl-games: A high-performance framework for reinforcement learning
Makoviichuk, D. and Makoviychuk, V · 2021
Cited alongside, same era.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
Visual dexterity: In-hand reorientation of novel and complex object shapes
Chen, T., Tippur, M., Wu, S., Kumar, V., Adelson, E., and Agrawal, P · 2023
Later among the works it cites.
Maniskill2: A unified benchmark for generalizable manipulation skills
Gu, J., Xiang, F., Li, X., Ling, Z., Liu, X., Mu, T., Tang, Y., Tao, S., Wei, X., Yao, Y., et al · 2023
Later among the works it cites.
Reward learning with intractable normalizing functions
Hoegerman, J. and Losey, D · 2023
Later among the works it cites.
Tree-planner: Efficient close-loop task planning with large language models
Hu, M., Mu, Y., Yu, X., Ding, M., Wu, S., Shao, W., Chen, Q., Wang, B., Qiao, Y., and Luo, P · 2023
Later among the works it cites.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aprel: A library for active preference-based reward learning algorithms
Bıyık, E., Talati, A., and Sadigh, D · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I · 2022
Cited alongside, same era.
Mehta, S. A. and Losey, D. P · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Language reward modulation for pretraining reinforcement learning
Adeniji, A., Xie, A., Sferrazza, C., Seo, Y., James, S., and Abbeel, P · 2023
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning and optimization
Bıyık, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D · 2023
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Later among the works it cites.
Reflect: Summarizing robot experiences for failure explanation and correction
Liu, Z., Bahety, A., and Song, S · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Later among the works it cites.
Robogen: Towards unleashing infinite data for automated robot learning via generative simulation
Wang, Y., Xian, Z., Chen, F., Wang, T.-H., Wang, Y., Fragkiadaki, K., Erickson, Z., Held, D., and Gan, C · 2023
Later among the works it cites.
Text2reward: Automated dense reward function generation for reinforcement learning
Xie, T., Zhao, S., Wu, C. H., Liu, Y., Luo, Q., Zhong, V., Yang, Y., and Yu, T · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Arenas, M. G., Chiang, H.-T. L., Erez, T., Hasenclever, L., Humplik, J., et al · 2023
Later among the works it cites.