Fetching the paper…
Reading the bibliography…
Reward engineering has long been a challenge in Reinforcement Learning (RL) research, as it often requires extensive human effort and iterative processes of trial-and-error to design effective reward functions.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Theory and application of reward shaping in reinforcement learning
Laud, A. D · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Earlier work this paper cites.
Variational inverse control with events: A general framework for data-driven reward definition
Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K · 2018
Cited alongside, same era.
Asking easy questions: A user-friendly approach to active reward learning
Biyik, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D · 2019
Cited alongside, same era.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning, 2019
OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., d. O. Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Cited alongside, same era.
Active preference-based gaussian process regression for reward learning
Biyik, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D · 2020
Cited alongside, same era.
Zero-shot reward specification via grounded natural language
Mahmoudieh, P., Pathak, D., and Darrell, T · 2022
Later among the works it cites.
Language reward modulation for pretraining reinforcement learning, 2023
Adeniji, A., Xie, A., Sferrazza, C., Seo, Y., James, S., and Abbeel, P · 2023
Later among the works it cites.
Towards generalizable zero-shot manipulation via translating human interaction plans
Bharadhwaj, H., Gupta, A., Kumar, V., and Tulsiani, S · 2023
Later among the works it cites.
Accelerating reinforcement learning of robotic manipulations via feedback from large language models
Chu, K., Zhao, X., Weber, C., Li, M., and Wermter, S · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Softgym: Benchmarking deep reinforcement learning for deformable object manipulation
Lin, X., Wang, Y., Olkin, J., and Held, D · 2021
Cited alongside, same era.
Learning multimodal rewards from rankings
Myers, V., Biyik, E., Anari, N., and Sadigh, D · 2021
Cited alongside, same era.
f-irl: Inverse reinforcement learning via state marginal matching
Ni, T., Sikchi, H., Wang, Y., Gupta, T., Lee, L., and Eysenbach, B · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Cited alongside, same era.
Human-to-robot imitation in the wild
Bahl, S., Gupta, A., and Pathak, D · 2022
Cited alongside, same era.
Later among the works it cites.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
Later among the works it cites.
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Later among the works it cites.
Lift: Unsupervised reinforcement learning with foundation models as teachers, 2023
Nam, T., Lee, J., Zhang, J., Hwang, S. J., Lim, J. J., and Pertsch, K · 2023
Later among the works it cites.
Nottingham, K., Ammanabrolu, P., Suhr, A., Choi, Y., Hajishirzi, H., Singh, S., and Fox, R · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Vision-language models are zero-shot reward models for reinforcement learning
Rocamonde, J., Montesinos, V., Nava, E., Perez, E., and Lindner, D · 2023
Later among the works it cites.
RoboCLIP: One demonstration is enough to learn robot policies
Sontakke, S. A., Zhang, J., Arnold, S., Pertsch, K., Biyik, E., Sadigh, D., Finn, C., and Itti, L · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Robogen: Towards unleashing infinite data for automated robot learning via generative simulation
Wang, Y., Xian, Z., Chen, F., Wang, T.-H., Wang, Y., Fragkiadaki, K., Erickson, Z., Held, D., and Gan, C · 2023
Later among the works it cites.
Text2reward: Automated dense reward function generation for reinforcement learning
Xie, T., Zhao, S., Wu, C. H., Liu, Y., Luo, Q., Zhong, V., Yang, Y., and Yu, T · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Gonzalez Arenas, M., Lewis Chiang, H.-T., Erez, T., Hasenclever, L., Humplik, J., Ichter, B., Xiao, T., Xu, P., Zeng, A., Zhang, T., Heess, N., Sadigh, D., Tan, J., Tassa, Y., and Xia, F · 2023
Later among the works it cites.