Fetching the paper…
Reading the bibliography…
In preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Gromov-wasserstein distances and the metric approach to object matching
Mémoli, F · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
An automated measure of mdp similarity for transfer in reinforcement learning
Ammar, H. B., Eaton, E., Taylor, M. E., Mocanu, D. C., Driessens, K., Weiss, G., and Tüyls, K · 2014
Earlier work this paper cites.
Tutorial on variational autoencoders
Doersch, C · 2016
Earlier work this paper cites.
Gromov-wasserstein averaging of kernel and distance matrices
Peyré, G., Cuturi, M., and Solomon, J · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Deepjdot: Deep joint distribution optimal transport for unsupervised domain adaptation
Damodaran, B. B., Kellenberger, B., Flamary, R., Tuia, D., and Courty, N · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Earlier work this paper cites.
Wasserstein distance guided representation learning for domain adaptation
Shen, J., Qu, Y., Zhang, W., and Yu, Y · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Earlier work this paper cites.
Optimal transport for structured data with application on graphs
Titouan, V., Courty, N., Tavenard, R., and Flamary, R · 2019
Earlier work this paper cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Earlier work this paper cites.
Gromov-wasserstein learning for graph matching and node embedding
Xu, H., Luo, D., Zha, H., and Duke, L. C · 2019
Cited alongside, same era.
Robust person re-identification by modelling feature uncertainty
Yu, T., Li, D., Yang, Y., Hospedales, T. M., and Xiang, T · 2019
Cited alongside, same era.
Data uncertainty learning in face recognition
Chang, J., Lan, Z., Cheng, C., and Wei, Y · 2020
Cited alongside, same era.
Graph optimal transport for cross-domain alignment
Chen, L., Gan, Z., Cheng, Y., Li, L., Carin, L., and Liu, J · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Gromov-wasserstein guided representation learning for cross-domain recommendation
Li, X., Qiu, Z., Zhao, X., Wang, Z., Zhang, Y., Xing, C., and Wu, X · 2022
Later among the works it cites.
Reward uncertainty for exploration in preference-based reinforcement learning
Liang, X., Shu, K., Lee, K., and Abbeel, P · 2022
Later among the works it cites.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Liu, R., Bai, F., Du, Y., and Yang, Y · 2022
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Imitation learning from pixel observations for continuous control
Cohen, S., Amos, B., Deisenroth, M. P., Henaff, M., Vinitsky, E., and Yarats, D · 2021
Cited alongside, same era.
Primal wasserstein imitation learning
Dadashi, R., Hussenot, L., Geist, M., and Pietquin, O · 2021
Cited alongside, same era.
Pot: Python optimal transport
Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Boisbunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., et al · 2021
Cited alongside, same era.
Probabilistic modeling of semantic ambiguity for scene graph generation
Yang, G., Zhang, J., Zhang, Y., Wu, B., and Yang, Y · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Cross-domain imitation learning via optimal transport
Fickinger, A., Cohen, S., Russell, S., and Amos, B · 2022
Cited alongside, same era.
Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
Park, J., Seo, Y., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2022
Later among the works it cites.
Picor: Multi-task deep reinforcement learning with policy correction
Bai, F., Zhang, H., Tao, T., Wu, Z., Wang, Y., and Xu, B · 2023
Closest in time.
Inverse preference learning: Preference-based rl without a reward function
Hejna, J. and Sadigh, D · 2023
Closest in time.
Few-shot preference learning for human-in-the-loop rl
Hejna III, D. J. and Sadigh, D · 2023
Closest in time.
Preference transformer: Modeling human preferences using transformers for rl
Kim, C., Park, J., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2023
Closest in time.
Optimal transport for offline imitation learning
Luo, Y., Jiang, Z., Cohen, S., Grefenstette, E., and Deisenroth, M. P · 2023
Closest in time.
Reinforcement learning from diverse human preferences
Xue, W., An, B., Yan, S., and Xu, Z · 2023
Closest in time.
Efficient preference-based reinforcement learning via aligned experience estimation
Bai, F., Zhao, R., Zhang, H., Cui, S., Wen, Y., Yang, Y., Xu, B., and Han, L · 2024
Closest in time.
SEABO: A simple search-based method for offline imitation learning
Lyu, J., Ma, X., Wan, L., Liu, R., Li, X., and Lu, Z · 2024
Closest in time.