Fetching the paper…
Reading the bibliography…
Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zero-shot performance on unseen tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019b · 1902
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019 · 1911
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Schema-guided dialogue state tracking task at dstc8
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020a · 2002
Earlier work this paper cites.
Coach: A coarse-to-fine approach for cross-domain slot filling
Zihan Liu, Genta Indra Winata, Peng Xu, and Pascale Fung. 2020 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Span-convert: Few-shot span extraction for dialog with pretrained conversational representations
Sam Coope, Tyler Farghly, Daniela Gerz, Ivan Vulić, and Matthew Henderson. 2020a · 2005
Earlier work this paper cites.
Understanding and improving information transfer in multi-task learning
Sen Wu, Hongyang R Zhang, and Christopher Ré. 2020b · 2005
Earlier work this paper cites.
Convex: Data-efficient and few-shot slot labeling
Matthew Henderson and Ivan Vulić. 2020 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederick P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Frames: a corpus for adding memory to goal-oriented dialogue systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Earlier work this paper cites.
Key-value retrieval networks for task-oriented dialogue
Mihail Eric, Lakshmi Krishnan, Francois Charette, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
End-to-end conversation modeling track in dstc6
Chiori Hori and Takaaki Hori. 2017 · 2017
Earlier work this paper cites.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017 · 2017
Earlier work this paper cites.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017 · 2017
Earlier work this paper cites.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Earlier work this paper cites.
QuAC: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018 · 2018
Earlier work this paper cites.
Abstractive dialogue summarization with sentence-gated modeling optimized by dialogue acts
Chih-Wen Goo and Yun-Nung Chen. 2018 · 2018
Earlier work this paper cites.
EmotionLines: An emotion corpus of multi-party conversations
Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang, and Lun-Wei Ku. 2018 · 2018
Earlier work this paper cites.
Microsoft dialogue challenge: Building end-to-end task-completion dialogue systems
Xiujun Li, Sarah Panda, Jingjing Liu, and Jianfeng Gao. 2018 · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Hello, it’s GPT-2 - how can I help you? towards the use of pretrained language models for task-oriented dialogue systems
Paweł Budzianowski and Ivan Vulić. 2019 · 2019
Earlier work this paper cites.
Taskmaster-1: Toward a realistic and diverse dialog dataset
Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, and Andy Cedilnik. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019a · 2019
Earlier work this paper cites.
Grounded response generation task at dstc7
Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019 · 2019
Earlier work this paper cites.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey Bigham. 2019 · 2019
Earlier work this paper cites.
An evaluation dataset for intent classification and out-of-scope prediction
Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019 · 2019
Earlier work this paper cites.
Pretraining methods for dialog context representation learning
Shikib Mehri, Evgeniia Razumovskaia, Tiancheng Zhao, and Maxine Eskenazi. 2019 · 2019
Earlier work this paper cites.
OpenDialKG: Explainable conversational reasoning with attention-based walks over knowledge graphs
Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Cited alongside, same era.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Dialfact: A benchmark for fact-checking in dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Conversations are not flat: Modeling the dynamic information flow across dialogue utterances
Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou. 2021 · 2021
Later among the works it cites.
Few-shot bot: Prompt-based learning for dialogue systems
Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Example-driven intent prediction with observers
Shikib Mehri and Mihail Eric. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019 · 2019
Cited alongside, same era.
Dialogue natural language inference
Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019 · 2019
Cited alongside, same era.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2020
Cited alongside, same era.
MuTual: A dataset for multi-turn dialogue reasoning
Leyang Cui, Yu Wu, Shujie Liu, Yue Zhang, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
GoEmotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020 · 2020
Cited alongside, same era.
doc2dial: A goal-oriented document-grounded dialogue dataset
Song Feng, Hui Wan, Chulaka Gunasekara, Siva Patel, Sachindra Joshi, and Luis Lastras. 2020a · 2020
Cited alongside, same era.
“none of the above”: Measure uncertainty in dialog response retrieval
Yulan Feng, Shikib Mehri, Maxine Eskenazi, and Tiancheng Zhao. 2020b · 2020
Cited alongside, same era.
Shikib Mehri and Maxine Eskenazi. 2021 · 2021
Later among the works it cites.
Self-training improves pre-training for few-shot learning in task-oriented dialog systems
Fei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai, Minlie Huang, and Boi Faltings. 2021b · 2021
Later among the works it cites.
Natural instructions: Benchmarking generalization to new tasks from natural language instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
I like fish, especially dolphins: Addressing contradictions in dialogue modeling
Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021 · 2021
Later among the works it cites.
Soloist: Building task bots at scale with transfer learning and machine teaching
Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
TIMEDIAL: Temporal commonsense reasoning in dialog
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui. 2021 · 2021
Later among the works it cites.
End-to-end learning of flowchart grounded task-oriented dialogs
Dinesh Raghu, Shantanu Agarwal, Sachindra Joshi, and Mausam. 2021 · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021 · 2021
Later among the works it cites.
Few-shot text generation with natural language instructions
Timo Schick and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
OTTers: One-turn topic transitions for open-domain dialogue
Karin Sevegnani, David M. Howcroft, Ioannis Konstas, and Verena Rieser. 2021 · 2021
Later among the works it cites.
Saferdialogues: Taking feedback gracefully after conversational safety failures
Megan Ung, Jing Xu, and Y-Lan Boureau. 2021 · 2021
Later among the works it cites.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick. 2021 · 2021
Later among the works it cites.
Do response selection models really know what’s next? utterance manipulation strategies for multi-turn response selection
Taesun Whang, Dongyub Lee, Dongsuk Oh, Chanhee Lee, Kijong Han, Dong-hun Lee, and Saebyeok Lee. 2021 · 2021
Later among the works it cites.
Qaconv: Question answering on informative conversations
Chien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021a · 2021
Later among the works it cites.
Ubar: Towards fully end-to-end task-oriented dialog system with gpt-2
Yunyi Yang, Yunhao Li, and Xiaojun Quan. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.
TuringAdvice: A generative and dynamic evaluation of language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi. 2021 · 2021
Later among the works it cites.
DynaEval: Unifying turn and dialogue level evaluation
Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021 · 2021
Later among the works it cites.
QMSum: A new benchmark for query-based multi-domain meeting summarization
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021a · 2021
Later among the works it cites.
Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections
Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. 2021b · 2021
Later among the works it cites.
Redwood: Using collision detection to grow a large-scale intent classification dataset
Stefan Larson and Kevin Leach. 2022 · 2022
Closest in time.
Unsupervised cross-task generalization via retrieval augmentation
Bill Yuchen Lin, Kangmin Tan, Chris Miller, Beiwen Tian, and Xiang Ren. 2022 · 2022
Closest in time.
Pretraining the noisy channel model for task-oriented dialogue
Qi Liu, Lei Yu, Laura Rimell, and Phil Blunsom. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Closest in time.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2022 · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Paul Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew Mingbo Dai, and Quoc V. Le. 2022 · 2022
Closest in time.
Balancing multi-domain corpora learning for open-domain response generation
Yujie Xing, Jinglun Cai, Nils Barlaug, Peng Liu, and Jon Atle Gulla. 2022 · 2022
Closest in time.
Zeroprompt: Scaling prompt-based pretraining to 1,000 tasks improves zero-shot generalization
Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang. 2022 · 2022
Closest in time.