Fetching the paper…
Reading the bibliography…
The availability of large pre-trained models is changing the landscape of Machine Learning research and practice, moving from a training-from-scratch to a fine-tuning paradigm.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 1910
Earlier work this paper cites.
Distributional Reinforcement Learning For Energy-Based Sequential Models
Tetiana Parshakova, Jean-Marc Andreoli, and Marc Dymetman · 1912
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard S. Sutton · 1984
Earlier work this paper cites.
Reinforcement-learning connectionist systems
Ronald J. Williams · 1987
Earlier work this paper cites.
Reinforcement comparison
Peter Dayan · 1990
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E. Hinton · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Empowerment: a universal agent-centric measure of control
A.S. Klyubin, D. Polani, and C.L. Nehaniv · 2005
Earlier work this paper cites.
A Tutorial on Energy-Based Learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
A unified energy-based framework for unsupervised learning
Marc’Aurelio Ranzato, Y-Lan Boureau, Sumit Chopra, and Yann LeCun · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano · 2009
Earlier work this paper cites.
Action and behavior: a free-energy formulation
Karl J Friston, Jean Daunizeau, James Kilner, and Stefan J Kiebel · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Importance Sampling
Art B. Owen · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Neural variational inference and learning in belief networks
Andriy Mnih and Karol Gregor · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Globally Normalized Transition-Based Neural Networks
Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins · 2016
Cited alongside, same era.
An Actor-Critic Algorithm for Sequence Prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Structured prediction energy networks
David Belanger and Andrew McCallum · 2016
Cited alongside, same era.
Neural text generation from structured data with application to the biography domain
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Later among the works it cites.
Controllable neural story plot generation via reward shaping
Pradyumna Tambwekar, Murtaza Dhuliawala, Lara J. Martin, Animesh Mehta, Brent Harrison, and Mark O. Riedl · 2019
Later among the works it cites.
A. Bakhtin, Y. Deng, S. Gross, Myle Ott, Marc’Aurelio Ranzato, and Arthur Szlam · 2020
Later among the works it cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Later among the works it cites.
Language gans falling short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rémi Lebret, David Grangier, and Michael Auli · 2016
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao · 2016
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Leshem Choshen, Lior Fox, Zohar Aizenbud, and Omri Abend · 2020
Later among the works it cites.
Pre-training transformers as energy-based cloze models
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Later among the works it cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2020
Later among the works it cites.
Residual energy-based models for text generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato · 2020
Later among the works it cites.
Action and perception as divergence minimization, 2020
Danijar Hafner, Pedro A. Ortega, Jimmy Ba, Thomas Parr, Karl Friston, and Nicolas Heess · 2020
Later among the works it cites.
Energy-based reranking: Improving neural machine translation using energy-based models
Subhajit Naskar, Pedram Rooshenas, Simeng Sun, Mohit Iyyer, and A. McCallum · 2020
Later among the works it cites.
Engine: Energy-based inference networks for non-autoregressive machine translation
Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, and Kevin Gimpel · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Hiroki Furuta, Tadashi Kozuno, Tatsuya Matsushima, Yutaka Matsuo, and Shixiang Shane Gu · 2021
Later among the works it cites.
Joint energy-based model training for better calibrated natural language understanding models
Tianxing He, Bryan McCann, Caiming Xiong, and Ehsan Hosseini-Asl · 2021
Later among the works it cites.
A distributional approach to controlled text generation
Muhammad Khalifa, Hady Elsahar, and Marc Dymetman · 2021
Later among the works it cites.
Revisiting the weaknesses of reinforcement learning for neural machine translation
Samuel Kiegeland and Julia Kreutzer · 2021
Later among the works it cites.
Energy-based models for code generation under compilability constraints
Tomasz Korbak, Hady Elsahar, Marc Dymetman, and Germán Kruszewski · 2021
Later among the works it cites.
Towards understanding and mitigating social biases in language models, 2021
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Later among the works it cites.
Understanding the origin of information-seeking exploration in probabilistic objectives for control, 2021
Beren Millidge, Alexander Tschantz, Anil Seth, and Christopher Buckley · 2021
Later among the works it cites.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Later among the works it cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Closest in time.
Red teaming language models with language models, 2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Closest in time.