Fetching the paper…
Reading the bibliography…
Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress.
Dynamic programming and markov processes
Ronald A Howard. 1960 · 1960
Earlier work this paper cites.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997 · 1997
Earlier work this paper cites.
Mixed-initiative interaction
James E Allen, Curry I Guinn, and Eric Horvitz. 1999 · 1999
Earlier work this paper cites.
How to build user simulators to train RL-based dialog systems
Weiyan Shi, Kun Qian, Xuewei Wang, and Zhou Yu. 2019 · 2000
Earlier work this paper cites.
Parallel monte-carlo tree search
Guillaume MJ B Chaslot, Mark HM Winands, and H Jaap van Den Herik. 2008 · 2008
Earlier work this paper cites.
Optimization and control
Richard Weber. 2010 · 2010
Earlier work this paper cites.
Multi-armed bandits with episode context
Christopher D Rosin. 2011 · 2011
Earlier work this paper cites.
Open loop search for general video game playing
Diego Perez Liebana, Jens Dieskau, Martin Hunermund, Sanaz Mostaghim, and Simon Lucas. 2015 · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 2017 · 2017
Earlier work this paper cites.
Multiwoz - a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Ultes Stefan, Ramadan Osman, and Milica Gašić. 2018 · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Earlier work this paper cites.
Zero-shot dialog generation with cross-domain latent actions
Tiancheng Zhao and Maxine Eskenazi. 2018 · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Persuasion for good: Towards a personalized persuasive dialogue system for social good
Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019 · 2019
Cited alongside, same era.
Adaptive dialog policy learning with hindsight and user modeling
Yan Cao, Keting Lu, Xiaoping Chen, and Shiqi Zhang. 2020 · 2020
Cited alongside, same era.
Bayes-adaptive monte-carlo planning and learning for goal-oriented dialogues
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim. 2020 · 2020
Cited alongside, same era.
Task-completion dialogue policy learning via Monte Carlo tree search with dueling network
Sihan Wang, Kaijie Zhou, Kunfeng Lai, and Jianping Shen. 2020 · 2020
Cited alongside, same era.
Galaxy: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, et al. 2022 · 2022
Later among the works it cites.
Soda: Million-scale dialogue distillation with social commonsense contextualization
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, et al. 2022 · 2022
Later among the works it cites.
Multi-stage prompting for knowledgeable dialogue generation
Zihan Liu, Mostofa Patwary, Ryan Prenger, Shrimai Prabhumoye, Wei Ping, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Later among the works it cites.
Openai: Introducing chatgpt
OpenAI. 2022 · 2022
Later among the works it cites.
Prompt learning for few-shot dialogue state tracking
Yuting Yang, Wenqiang Lei, Juan Cao, Jintao Li, and Tat-Seng Chua. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A hierarchy mcts algorithm for the automated pcb routing
Cong Zhang, Huilin Jin, Jienan Chen, Jinkuan Zhu, and Jinting Luo. 2020a · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Cited alongside, same era.
Legoeval: An open-source toolkit for dialogue system evaluation via crowdsourcing
Yu Li, Josh Arnold, Feifan Yan, Weiyan Shi, and Zhou Yu. 2021 · 2021
Cited alongside, same era.
Towards emotional support dialog systems
Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Few-shot bot: Prompt-based learning for dialogue systems
Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021 · 2021
Cited alongside, same era.
Schema-guided paradigm for zero-shot dialog
Shikib Mehri and Maxine Eskenazi. 2021 · 2021
Cited alongside, same era.
Want to reduce labeling cost? gpt-3 can help
Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
PLACES: Prompting language models for social conversation synthesis
Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andy Rosenbaum, Yang Liu, Zhou Yu, and Dilek Hakkani-Tur. 2023a · 2023
Closest in time.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Closest in time.
Annollm: Making large language models to be better crowdsourced annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, A Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023 · 2023
Closest in time.
Commonsense-aware prompting for controllable empathetic dialogue generation
Yiren Liu and Halil Kilicoglu. 2023 · 2023
Closest in time.
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Conversational tree search: A new hybrid dialog task
Dirk Väth, Lindsey Vanderlyn, and Ngoc Thang Vu. 2023 · 2023
Closest in time.
How far can camels go? exploring the state of instruction tuning on open resources
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
Qiang Zhang, Jason Naradowsky, and Yusuke Miyao. 2023 · 2023
Closest in time.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gökhan Tür, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018 · 2069
Closest in time.