Fetching the paper…
Reading the bibliography…
Preference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H., Misra, N., Hnizdo, V., Fedorowicz, A., and Demchuk, E · 2003
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Grandvalet, Y. and Bengio, Y · 2004
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
Preference-based policy iteration: Leveraging preference learning for reinforcement learning
Cheng, W., Fürnkranz, J., Hüllermeier, E., and Park, S.-H · 2011
Earlier work this paper cites.
Training deep neural-networks using a noise adaptation layer
Goldberger, J. and Ben-Reuven, E · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Robust loss functions under label noise for deep neural networks
Ghosh, A., Kumar, H., and Sastry, P. S · 2017
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhang, Z. and Sabuncu, M · 2018
Earlier work this paper cites.
Curriculum loss: Robust learning and generalization against label corruption
Lyu, Y. and Tsang, I. W · 2019
Earlier work this paper cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z · 2020
Cited alongside, same era.
Weakly supervised learning with side information for noisy labeled images
Cheng, L., Zhou, X., Zhao, L., Li, D., Shang, H., Zheng, Y., Pan, P., and Xu, Y · 2020
Cited alongside, same era.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Li, M., Soltanolkotabi, M., and Oymak, S · 2020
Cited alongside, same era.
Does label smoothing mitigate label noise?
Lukasik, M., Bhojanapalli, S., Menon, A., and Kumar, S · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Trochim, P., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., et al · 2020
Cited alongside, same era.
To smooth or not? when label smoothing meets noisy labels
Wei, J., Liu, H., Liu, T., Niu, G., Sugiyama, M., and Liu, Y · 2021
Later among the works it cites.
Reincarnating reinforcement learning: Reusing prior computation to accelerate progress
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2022
Later among the works it cites.
Towards human-level bimanual dexterous manipulation with reinforcement learning
Chen, Y., Wu, T., Wang, S., Feng, X., Jiang, J., Lu, Z., McAleer, S., Dong, H., Zhu, S.-C., and Yang, Y · 2022
Later among the works it cites.
Preference transformer: Modeling human preferences using transformers for rl
Kim, C., Park, J., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2022
Later among the works it cites.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Liu, R., Bai, F., Du, Y., and Yang, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust early-learning: Hindering the memorization of noisy labels
Xia, X., Liu, T., Han, B., Gong, C., Wang, N., Ge, Z., and Chang, Y · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving
Zhou, M., Luo, J., Villella, J., Yang, Y., Rusu, D., Miao, J., Zhang, W., Alban, M., Fadakar, I., Chen, Z., et al · 2020
Cited alongside, same era.
Can cross entropy loss be robust to label noise?
Feng, L., Shu, S., Lin, Z., Lv, F., Li, L., and An, B · 2021
Cited alongside, same era.
Reward uncertainty for exploration in preference-based reinforcement learning
Liang, X., Shu, K., Lee, K., and Abbeel, P · 2021
Cited alongside, same era.
Behavior from the void: Unsupervised active pre-training
Liu, H. and Abbeel, P · 2021
Cited alongside, same era.
Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
Park, J., Seo, Y., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Perolat, J., De Vylder, B., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., et al · 2022
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Song, H., Kim, M., Park, D., Shin, Y., and Lee, J.-G · 2022
Later among the works it cites.
Few-shot preference learning for human-in-the-loop rl
Hejna III, D. J. and Sadigh, D · 2023
Later among the works it cites.
Champion-level drone racing using deep reinforcement learning
Kaufmann, E., Bauersfeld, L., Loquercio, A., Müller, M., Koltun, V., and Scaramuzza, D · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Later among the works it cites.
Jump-start reinforcement learning
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., et al · 2023
Later among the works it cites.
Reinforcement learning from diverse human preferences
Xue, W., An, B., Yan, S., and Xu, Z · 2023
Later among the works it cites.
Sc-tune: Unleashing self-consistent referential comprehension in large vision language models
Yue, T., Cheng, J., Guo, L., Dai, X., Zhao, Z., He, X., Xiong, G., Lv, Y., and Liu, J · 2024
Closest in time.