Fetching the paper…
Reading the bibliography…
Large language models distill broad knowledge from text corpora.
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 1910
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Learning to Forget: Continual Prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins · 2000
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2004
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Morel : Model-based offline reinforcement learning, 2020
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2005
Earlier work this paper cites.
Temporal abstraction in temporal-difference networks
Eddie Rafols, Anna Koop, and Richard S Sutton · 2005
Earlier work this paper cites.
Mopo: Model-based offline policy optimization, 2020
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2005
Earlier work this paper cites.
Awac: Accelerating online reinforcement learning with offline datasets, 2020
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2006
Earlier work this paper cites.
Critic regularized regression, 2020
Ziyu Wang, Alexander Novikov, Konrad Zolna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2006
Earlier work this paper cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl, 2020
Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu · 2007
Earlier work this paper cites.
Text generation by learning from demonstrations, 2020
Richard Yuanzhe Pang and He He · 2009
Earlier work this paper cites.
Human-centric dialog training via offline reinforcement learning, 2020
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Shane Gu, and Rosalind Picard · 2010
Earlier work this paper cites.
Sequence level training with recurrent neural networks, 2015
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Earlier work this paper cites.
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Learning to translate in real-time with neural machine translation, 2016
Jiatao Gu, Graham Neubig, Kyunghyun Cho, and Victor O. K. Li · 2016
Earlier work this paper cites.
On-line active reward learning for policy optimisation in spoken dialogue systems, 2016
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young · 2016
Cited alongside, same era.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Cited alongside, same era.
A review of evaluation techniques for social dialogue systems
Amanda Cercas Curry, Helen Hastie, and Verena Rieser · 2017
Cited alongside, same era.
Learning cooperative visual dialog agents with deep reinforcement learning, 2017
Abhishek Das, Satwik Kottur, José M. F. Moura, Stefan Lee, and Dhruv Batra · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Offline rl without off-policy evaluation, 2021
David Brandfonbrener, William F. Whitney, Rajesh Ranganath, and Joan Bruna · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling, 2021
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hafez: an interactive poetry generation system
Marjan Ghazvininejad, Xing Shi, Jay Priyadarshi, and Kevin Knight · 2017
Cited alongside, same era.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernandez-Lobato, R. E. Turner, and D. Eck · 2017
Cited alongside, same era.
Learning to decode for future success, 2017
Jiwei Li, Will Monroe, and Dan Jurafsky · 2017
Cited alongside, same era.
A deep reinforced model for abstractive summarization, 2017
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Recent trends in deep learning based natural language processing, 2017
Tom Young, Devamanyu Hazarika, Soujanya Poria, and Erik Cambria · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak · 2021
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?, 2021
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Towards automatic evaluation of dialog systems: A model-free off-policy evaluation approach
Haoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao, and Wei Wei · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning, 2021
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Later among the works it cites.
GeDi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani · 2021
Later among the works it cites.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2021
Later among the works it cites.
Noisy channel language model prompting for few-shot text classification, 2021
Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2021
Later among the works it cites.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Later among the works it cites.
GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim · 2022
Closest in time.
Daniel Lokshtanov and Bernardo Subercaseaux · 2022
Closest in time.
Context-aware language modeling for goal-oriented dialogue systems, 2022
Charlie Snell, Sherry Yang, Justin Fu, Yi Su, and Sergey Levine · 2022
Closest in time.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning, 2022
Siddharth Verma, Justin Fu, Mengjiao Yang, and Sergey Levine · 2022
Closest in time.