Improved optimistic algorithms for logistic bandits
L. Faury, M. Abeille, C. Calauzènes, and O. Fercoq · 2020
Later among the works it cites.
Conservative Q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Later among the works it cites.
Preference-based reinforcement learning with finite-time guarantees
Y. Xu, R. Wang, L. Yang, A. Singh, and A. Dubrawski · 2020
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2021
Later among the works it cites.
Offline reinforcement learning with implicit Q-learning
Original
I. Kostrikov, A. Nair, and S. Levine · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Original
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Later among the works it cites.
Dueling rl: reinforcement learning with trajectory preferences
Original
A. Pacchiano, A. Saha, and J. Lee · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Original
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
A. Zanette, M. J. Wainwright, and E. Brunskill · 2021
Later among the works it cites.
Advances in preference-based reinforcement learning: A review
Y. Abdelkareem, S. Shehata, and F. Karray · 2022
Later among the works it cites.
On the theory of reinforcement learning with once-per-episode feedback, 2022
N. S. Chatterji, A. Pacchiano, P. L. Bartlett, and M. I. Jordan · 2022
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
X. Chen, H. Zhong, Z. Yang, Z. Wang, and L. Wang · 2022
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Original
C.-A. Cheng, T. Xie, N. Jiang, and A. Agarwal · 2022
Later among the works it cites.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Original
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization
Original
L. Gao, J. Schulman, and J. Hilton · 2022
Later among the works it cites.
Exploiting correlation to achieve faster learning rates in low-rank preference bandits
S. Ghoshal and A. Saha · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Original
A. Glaese, N. McAleese, M. Tr e ⋅ \underset{\cdot}{e} bacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Later among the works it cites.
Pessimism for offline linear contextual bandits using ℓ p \ell_{p} confidence sets
Original
G. Li, C. Ma, and N. Srebro · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Original
J. Menick, M. Trebacz, V. Mikulik, J. Aslanides, F. Song, M. Chadwick, M. Glaese, S. Young, L. Campbell-Gillingham, G. Irving, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization
Original
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2022
Later among the works it cites.
Efficient and optimal algorithms for contextual dueling bandits under realizability
A. Saha and A. Krishnamurthy · 2022
Later among the works it cites.
Provably efficient offline reinforcement learning with trajectory-wise reward
Original
T. Xu and Y. Liang · 2022
Later among the works it cites.
When is realizability sufficient for off-policy reinforcement learning?
Original
A. Zanette · 2022
Later among the works it cites.
Benchmarks and algorithms for offline preference-based reward learning
Original
D. Shin, A. D. Dragan, and D. S. Brown · 2023
Closest in time.