Fetching the paper…

COPR: Continual Learning Human Preference through Optimal Policy Regularization · Around