Fetching the paper…

COPR: Continual Human Preference Learning via Optimal Policy Regularization · Around