Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has demonstrated compelling performance in robotic tasks, but its success often hinges on the design of complex, ad hoc reward functions.
Telling more than we can know: Verbal reports on mental processes
Nisbett, R. E. and Wilson, T. D · 1977
Earlier work this paper cites.
Verbal reports as data
Ericsson, K. A. and Simon, H. A · 1980
Earlier work this paper cites.
Social foundations of thought and action: A social-cognitive view, 1987
Locke, E. A · 1987
Earlier work this paper cites.
Eliciting knowledge from experts: A methodological analysis
Hoffman, R. R., Shadbolt, N. R., Burton, A. M., and Klein, G · 1995
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Reward function and initial values: Better choices for accelerated goal-directed reinforcement learning
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2006
Earlier work this paper cites.
The implications of research on expertise for curriculum and pedagogy
Feldon, D. F · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
Robot Learning from Human Teachers
Chernova, S. and Thomaz, A. L · 2014
Earlier work this paper cites.
Learning by thinking: How reflection improves performance
Di Stefano, G., Gino, F., Pisano, G., and Staats, B · 2014
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Learning from demonstration for shaping through inverse reinforcement learning
Suay, H. B., Brys, T., Taylor, M. E., and Chernova, S · 2016
Cited alongside, same era.
Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation
Tai, L., Paolo, G., and Liu, M · 2017
Cited alongside, same era.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., and Farhadi, A · 2017
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2018
Cited alongside, same era.
Robot learning from demonstration in robotic assembly: A survey
The perils of trial-and-error reward design: Misdesign through overfitting and invalid task specifications
Booth, S., Knox, W. B., Shah, J., Niekum, S., Stone, P., and Allievi, A · 2023
Later among the works it cites.
Vision-language models as success detectors
Du, Y., Konyushkova, K., Denil, M., Raju, A., Landon, J., Hill, F., de Freitas, N., and Cabi, S · 2023
Later among the works it cites.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Evaluating a large language model’s ability to solve programming exercises from an introductory bioinformatics course
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, Z. and Hu, H · 2018
Cited alongside, same era.
Joint goal and strategy inference across heterogeneous demonstrators via reward network distillation
Chen, L., Paleja, R., Ghuy, M., and Gombolay, M · 2020
Cited alongside, same era.
Recent advances in robot learning from demonstration
Ravichandar, H., Polydoros, A. S., Chernova, S., and Billard, A · 2020
Cited alongside, same era.
A survey of inverse reinforcement learning: Challenges, methods and progress
Arora, S. and Doshi, P · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al · 2021
Cited alongside, same era.
Fast lifelong adaptive inverse reinforcement learning from demonstrations
Chen, L., Jayanthi, S., Paleja, R. R., Martin, D., Zakharov, V., and Gombolay, M · 2022
Cited alongside, same era.
Piccolo, S. R., Denny, P., Luxton-Reilly, A., Payne, S. H., and Ridge, P. G · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Prompt a robot to walk with large language models
Wang, Y.-J., Zhang, B., Chen, J., and Sreenath, K · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Gonzalez Arenas, M., Lewis Chiang, H.-T., Erez, T., Hasenclever, L., Humplik, J., Ichter, B., Xiao, T., Xu, P., Zeng, A., Zhang, T., Heess, N., Sadigh, D., Tan, J., Tassa, Y., and Xia, F · 2023
Later among the works it cites.
A review of reward functions for reinforcement learning in the context of autonomous driving
Abouelazm, A., Michel, J., and Zoellner, J. M · 2024
Closest in time.
Interactive human-robot teaching recovers and builds trust, even with imperfect learners
Chi, V. B. and Malle, B. F · 2024
Closest in time.
Behavior alignment via reward function optimization
Gupta, D., Chandak, Y., Jordan, S., Thomas, P. S., and C da Silva, B · 2024
Closest in time.
Generative expressive robot behaviors using large language models
Mahadevan, K., Chien, J., Brown, N., Xu, Z., Parada, C., Xia, F., Zeng, A., Takayama, L., and Sadigh, D · 2024
Closest in time.
Roboclip: One demonstration is enough to learn robot policies
Sontakke, S., Zhang, J., Arnold, S., Pertsch, K., Bıyık, E., Sadigh, D., Finn, C., and Itti, L · 2024
Closest in time.
Rl-vlm-f: Reinforcement learning from vision language foundation model feedback
Wang, Y., Sun, Z., Zhang, J., Xian, Z., Biyik, E., Held, D., and Erickson, Z · 2024
Closest in time.