Fetching the paper…

Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation · Around