Fetching the paper…

Fine-tuning Language Models with Generative Adversarial Reward Modelling · Around