Fetching the paper…

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models · Around