Fetching the paper…
Reading the bibliography…
To align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Is deep reinforcement learning really superhuman on Atari? leveling the playing field
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde. 2019 · 1908
Earlier work this paper cites.
Combining feature and instance attribution to detect artifacts
Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron Wallace. 2022 · 1946
Earlier work this paper cites.
Problems of monetary management: the UK experience in papers in monetary economics
Charles Goodhart. 1975 · 1975
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
Ferenc Huszár. 2015 · 2015
Earlier work this paper cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Earlier work this paper cites.
What to do about non-standard (or non-canonical) language in NLP
Barbara Plank. 2016 · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Earlier work this paper cites.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Earlier work this paper cites.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Long text generation via adversarial training with leaked information
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang. 2018 · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018 · 2018
Earlier work this paper cites.
Learning to run challenge solutions: Adapting reinforcement learning methods for neuromusculoskeletal environments
Łukasz Kidziński, Sharada Prasanna Mohanty, Carmichael F Ong, Zhewei Huang, Shuchang Zhou, Anton Pechenko, Adam Stelmaszczyk, Piotr Jarosik, Mikhail Pavlov, Sergey Kolesnikov, et al. 2018 · 2018
Earlier work this paper cites.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Multi-reward reinforced summarization with saliency and entailment
Ramakanth Pasunuru and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Generalization in generation: A closer look at exposure bias
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving. 2021 · 2021
Later among the works it cites.
A distributional approach to controlled text generation
Muhammad Khalifa, Hady Elsahar, and Marc Dymetman. 2021 · 2021
Later among the works it cites.
Revisiting the weaknesses of reinforcement learning for neural machine translation
Samuel Kiegeland and Julia Kreutzer. 2021 · 2021
Later among the works it cites.
Text generation by learning from demonstrations
Richard Yuanzhe Pang and He He. 2021 · 2021
Later among the works it cites.
AgreeSum: Agreement-oriented multi-document summarization
Richard Yuanzhe Pang, Adam Lelkes, Vinh Tran, and Cong Yu. 2021 · 2021
Later among the works it cites.
Teach me to explain: A review of datasets for explainable natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Florian Schmidt. 2019 · 2019
Cited alongside, same era.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne. 2019 · 2019
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch. 2020 · 2020
Cited alongside, same era.
Findings of the 2020 conference on machine translation (WMT20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Cited alongside, same era.
Human-paraphrased references improve neural machine translation
Markus Freitag, George Foster, David Grangier, and Colin Cherry. 2020 · 2020
Cited alongside, same era.
Q-learning with language model for edit-based unsupervised summarization
Ryosuke Kohita, Akifumi Wachi, Yang Zhao, and Ryuki Tachibana. 2020 · 2020
Cited alongside, same era.
Specification gaming: the flip side of AI ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg. 2020 · 2020
Cited alongside, same era.
Sarah Wiegreffe and Ana Marasovic. 2021 · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021 · 2021
Later among the works it cites.
The cringe loss: Learning what language not to model
Leonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, and Jason Weston. 2022 · 2022
Closest in time.
Identifying weaknesses in machine translation metrics through minimum Bayes risk decoding: A case study for COMET
Chantal Amrhein and Rico Sennrich. 2022 · 2022
Closest in time.
Why exposure bias matters: An imitation learning perspective of error accumulation in language generation
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Cheung. 2022 · 2022
Closest in time.
Nano: Nested human-in-the-loop reward learning for few-shot language model control
Xiang Fan, Yiwei Lyu, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Closest in time.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton. 2022 · 2022
Closest in time.
Objective robustness in deep reinforcement learning
Jack Koch, Lauro Langosco, Jacob Pfau, James Le, and Lee Sharkey. 2022 · 2022
Closest in time.
RL with KL penalties is better viewed as bayesian inference
Tomasz Korbak, Ethan Perez, and Christopher L Buckley. 2022 · 2022
Closest in time.
Characterizing and addressing the issue of oversmoothing in neural autoregressive sequence modeling
Ilia Kulikov, Maksim Eremeev, and Kyunghyun Cho. 2022 · 2022
Closest in time.
Comparing BERT-based reward functions for deep reinforcement learning in machine translation
Yuki Nakatani, Tomoyuki Kajiwara, and Takashi Ninomiya. 2022 · 2022
Closest in time.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt. 2022 · 2022
Closest in time.
Amortized noisy channel neural machine translation
Richard Yuanzhe Pang, He He, and Kyunghyun Cho. 2022 · 2022
Closest in time.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. 2022 · 2022
Closest in time.
Defining and characterizing reward hacking
Joar Skalse, Nikolaus HR Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Closest in time.
Dual generator offline reinforcement learning
Quan Vuong, Aviral Kumar, Sergey Levine, and Yevgen Chebotar. 2022 · 2022
Closest in time.
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover. 2022 · 2022
Closest in time.
Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.