A closer look at invalid action masking in policy gradient algorithms
Original
Shengyi Huang and Santiago Ontañón · 2020
Later among the works it cites.
Human-centric dialog training via offline reinforcement learning
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2020
Later among the works it cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Later among the works it cites.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 2020
Later among the works it cites.
ToTTo: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Later among the works it cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Later among the works it cites.
Language Learning in Interactive Environments
Prithviraj Ammanabrolu · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
Relating neural text degeneration to exposure bias
Ting-Rui Chiang and Yun-Nung Chen · 2021
Later among the works it cites.
Revisiting the weaknesses of reinforcement learning for neural machine translation
Samuel Kiegeland and Julia Kreutzer · 2021
Later among the works it cites.
Offline reinforcement learning from human feedback in real-world sequence-to-sequence tasks
Julia Kreutzer, Stefan Riezler, and Carolin Lawrence · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Original
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Text generation by learning from demonstrations
Richard Yuanzhe Pang and He He · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap
Gokul Swamy, Sanjiban Choudhury, J Andrew Bagnell, and Steven Wu · 2021
Later among the works it cites.
Textgail: Generative adversarial imitation learning for text generation
Qingyang Wu, Lei Li, and Zhou Yu · 2021
Later among the works it cites.
Aligning to social norms and values in interactive narratives
Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi, and Yejin Choi · 2022
Closest in time.
Why exposure bias matters: An imitation learning perspective of error accumulation in language generation
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Cheung · 2022
Closest in time.
Consistent dropout for policy gradient reinforcement learning
Original
Matthew Hausknecht and Nolan Wagener · 2022
Closest in time.
SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst · 2022
Closest in time.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Liwei Jiang, Jack Hessel, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Closest in time.
Learning natural language generation with truncated reinforcement learning
Alice Martin, Guillaume Quispe, Charles Ollion, Sylvain Le Corff, Florian Strub, and Olivier Pietquin · 2022
Closest in time.
Training language models to follow instructions with human feedback
Original
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Closest in time.
Offline rl for natural language generation with implicit language q learning
Original
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine · 2022
Closest in time.