Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are powerful dialogue agents, but specializing them towards fulfilling a specific function can be challenging.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro. 1994 · 1994
Earlier work this paper cites.
Why did td-gammon work?
Jordan Pollack and Alan Blair. 1996 · 1996
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Deploying lifelong open-domain dialogue learning
Kurt Shuster, Jack Urbanek, Emily Dinan, Arthur Szlam, and Jason Weston. 2020 · 2008
Earlier work this paper cites.
Reinforcement learning in the game of othello: Learning against a fixed opponent and learning from self-play
Michiel Van Der Ree and Marco Wiering. 2013 · 2013
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016 · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017 · 2017
Earlier work this paper cites.
An optimal transportation approach for assessing almost stochastic order
Eustasio Del Barrio, Juan A Cuesta-Albertos, and Carlos Matrán. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018 · 2018
Earlier work this paper cites.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gökhan Tür. 2018a · 2018
Earlier work this paper cites.
Deep dominance - how to properly compare deep neural models
Rotem Dror, Segev Shlomov, and Roi Reichart. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Learning to speak and act in a fantasy text adventure game
Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al. 2020 · 2020
Earlier work this paper cites.
Action-based conversations dataset: A corpus for building more in-depth task-oriented dialogue systems
Derek Chen, Howard Chen, Yi Yang, Alexander Lin, and Zhou Yu. 2021 · 2021
Earlier work this paper cites.
A survey on bias in deep nlp
Ismael Garrido-Muñoz, Arturo Montejo-Ráez, Fernando Martínez-Santiago, and L Alfonso Ureña-López. 2021 · 2021
Earlier work this paper cites.
Deberta: decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021b · 2021
Cited alongside, same era.
A survey on gender bias in natural language processing
Karolina Stanczak and Isabelle Augenstein. 2021 · 2021
Cited alongside, same era.
Disembodied machine learning: On the illusion of objectivity in nlp
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein. 2021 · 2021
Cited alongside, same era.
Gpt3mix: Leveraging large-scale language models for text augmentation
Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woo-Myoung Park. 2021 · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 · 2022
Chataug: Leveraging chatgpt for text data augmentation
Haixing Dai, Zhengliang Liu, Wenxiong Liao, Xiaoke Huang, Zihao Wu, Lin Zhao, Wei Liu, Ninghao Liu, Sheng Li, Dajiang Zhu, et al. 2023 · 2023
Later among the works it cites.
Openllama: An open reproduction of llama
Xinyang Geng and Hao Liu. 2023 · 2023
Later among the works it cites.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon. 2023 · 2023
Later among the works it cites.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas. 2023 · 2023
Later among the works it cites.
Large language models can be used to effectively scale spear phishing campaigns
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Controllable dialogue simulation with in-context learning
Zekun Li, Wenhu Chen, Shiyang Li, Hong Wang, Jing Qian, and Xifeng Yan. 2022 · 2022
Cited alongside, same era.
Hong Liu, Yucheng Cai, Zhijian Ou, Yi Huang, and Junlan Feng. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Neural theory-of-mind? on the limits of social intelligence in large lms
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022 · 2022
Cited alongside, same era.
deep-significance: Easy and meaningful signifcance testing in the age of neural networks
Dennis Ulmer, Christian Hardmeier, and Jes Frellsen. 2022 · 2022
Cited alongside, same era.
Julian Hazell. 2023 · 2023
Later among the works it cites.
Prompt-based methods may underestimate large language models’ linguistic generalizations
Jennifer Hu and Roger Levy. 2023 · 2023
Later among the works it cites.
Training socially aligned language models in simulated human society
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M Dai, Diyi Yang, and Soroush Vosoughi. 2023 · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML NLP Team. 2023 · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Later among the works it cites.
Refiner: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. 2023 · 2023
Later among the works it cites.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2023 · 2023
Later among the works it cites.
Model dementia: Generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2023 · 2023
Later among the works it cites.
Redpajama-data: An open source recipe to reproduce llama training dataset
Together Computer. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Sgp-tod: Building task bots effortlessly via schema-guided llm prompting
Xiaoying Zhang, Baolin Peng, Kun Li, Jingyan Zhou, and Helen Meng. 2023 · 2023
Later among the works it cites.