Fetching the paper…
Reading the bibliography…
Dialogue Policy Learning is a key component in a task-oriented dialogue system (TDS) that decides the next action of the system given the dialogue state at each turn.
Bert for joint intent classification and slot filling
Qian Chen, Zhu Zhuo, and Wen Wang. 2019a · 1902
Earlier work this paper cites.
Tiancheng Zhao, Kaige Xie, and Maxine Eskenazi. 2019 · 1902
Earlier work this paper cites.
Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention
Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, and William Yang Wang. 2019b · 1905
Earlier work this paper cites.
Meta-Learning for Low-resource Natural Language Generation in Task-oriented Dialogue Systems
Fei Mi, Minlie Huang, Jiyong Zhang, and Boi Faltings. 2019 · 1905
Earlier work this paper cites.
Transferable multi-domain state generator for task-oriented dialogue systems
Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, and Pascale Fung. 2019 · 1905
Earlier work this paper cites.
Budgeted Policy Learning for Task-Oriented Dialogue Systems
Zhirui Zhang, Xiujun Li, Jianfeng Gao, and Enhong Chen. 2019 · 1906
Earlier work this paper cites.
Dialog State Tracking: A Neural Reading Comprehension Approach
Shuyang Gao, Abhishek Sethi, Sanchit Agarwal, Tagyoung Chung, and Dilek Hakkani-Tur. 2019 · 1908
Earlier work this paper cites.
Modeling multi-action policy for task-oriented dialogues
Lei Shu, Hu Xu, Bing Liu, and Piero Molino. 2019 · 1908
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton. 1992 · 1992
Earlier work this paper cites.
User modeling for spoken dialogue system evaluation
Wieland Eckert, Esther Levin, and Roberto Pieraccini. 1997 · 1997
Earlier work this paper cites.
PARADISE: A framework for evaluating spoken dialogue agents
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart Russell. 1998 · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Ng. 1999 · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh. 1999 · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
A stochastic model of human-machine interaction for learning dialog strategies
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000 · 2000
Earlier work this paper cites.
Reinforcement Learning for Spoken Dialogue Systems
Satinder Singh, Michael Kearns, Diane Litman, and Marilyn Walker. 2000 · 2000
Earlier work this paper cites.
An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email
Marilyn A Walker. 2000 · 2000
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker. 2002 · 2002
Earlier work this paper cites.
Recent Advances and Challenges in Task-oriented Dialog System
Zheng Zhang, Ryuichi Takanobu, Qi Zhu, Minlie Huang, and Xiaoyan Zhu. 2020b · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y. Ng. 2004 · 2004
Earlier work this paper cites.
Learning Dialog Policies from Weak Demonstrations
Gabriel Gordon-Hall, Philip John Gorinski, and Shay B. Cohen. 2020a · 2004
Earlier work this paper cites.
Show Us the Way: Learning to Manage Dialog from Demonstrations
Gabriel Gordon-Hall, Philip John Gorinski, Gerasimos Lampouras, and Ignacio Iacobacci. 2020b · 2004
Earlier work this paper cites.
Guided Dialog Policy Learning without Adversarial Learning in the Loop
Ziming Li, Sungjin Lee, Baolin Peng, Jinchao Li, Julia Kiseleva, Maarten de Rijke, Shahin Shayandeh, and Jianfeng Gao. 2020b · 2004
Earlier work this paper cites.
AutoEG: Automated Experience Grafting for Off-Policy Deep Reinforcement Learning
Keting Lu, Shiqi Zhang, and Xiaoping Chen. 2020 · 2004
Earlier work this paper cites.
Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward Decomposition
Ryuichi Takanobu, Runze Liang, and Minlie Huang. 2020a · 2004
Earlier work this paper cites.
Multi-domain dialogue acts and response co-generation
Kai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan, and Jianxing Yu. 2020b · 2004
Earlier work this paper cites.
Learning Goal-oriented Dialogue Policy with Opposite Agent Awareness
Zheng Zhang, Lizi Liao, Xiaoyan Zhu, Tat-Seng Chua, Zitao Liu, Yan Huang, and Minlie Huang. 2020a · 2004
Earlier work this paper cites.
Adaptive Dialog Policy Learning with Hindsight and User Modeling
Yan Cao, Keting Lu, Xiaoping Chen, and Shiqi Zhang. 2020 · 2005
Earlier work this paper cites.
A Survey on Dialog Management: Recent Advances and Challenges
Yinpei Dai, Huihua Yu, Yixuan Jiang, Chengguang Tang, Yongbin Li, and Jian Sun. 2020 · 2005
Earlier work this paper cites.
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation
Xinting Huang, Jianzhong Qi, Yu Sun, and Rui Zhang. 2020 · 2005
Earlier work this paper cites.
Ryuichi Takanobu, Qi Zhu, Jinchao Li, Baolin Peng, Jianfeng Gao, and Minlie Huang. 2020b · 2005
Earlier work this paper cites.
Autonomous helicopter flight via reinforcement learning
Andrew Y Ng, H Jin Kim, Michael I Jordan, and Shankar Sastry. 2006 · 2006
Earlier work this paper cites.
Yumo Xu, Chenguang Zhu, Baolin Peng, and Michael Zeng. 2020 · 2006
Cited alongside, same era.
Spoken language understanding: a survey
Renato De Mori. 2007 · 2007
Cited alongside, same era.
Creating spoken dialogue characters from corpora without annotations
Sudeep Gandhe and David Traum. 2007 · 2007
Cited alongside, same era.
Agenda-Based User Simulation for Bootstrapping a POMDP Dialogue System
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Cited alongside, same era.
Assessing dialog system user simulation evaluation measures using human judges
Hua Ai and Diane Litman. 2008 · 2008
Cited alongside, same era.
Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets
Speaker Role Contextual Modeling for Language Understanding and Dialogue Policy Learning
Ta-Chung Chi, Po-Chun Chen, Shang-Yu Su, and Yun-Nung Chen. 2017 · 2017
Later among the works it cites.
A copy-augmented sequence-to-sequence architecture gives good performance on task-oriented dialogue
Mihail Eric and Christopher D Manning. 2017 · 2017
Later among the works it cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Later among the works it cites.
Deal or no deal? end-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Henderson, Oliver Lemon, and Kallirroi Georgila. 2008 · 2008
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal. 2008 · 2008
Cited alongside, same era.
Evaluating user simulations with the cramér–von mises divergence
Jason D Williams. 2008 · 2008
Cited alongside, same era.
The hidden agenda user simulation model
Jost Schatzmann and Steve Young. 2009 · 2009
Cited alongside, same era.
Learning the Reward Model of Dialogue POMDPs from Data
Abdeslam Boularias, Hamid R Chinaei, and Brahim Chaib-draa. 2010 · 2010
Cited alongside, same era.
User simulation in dialogue systems using inverse reinforcement learning
Senthilkumar Chandramohan, Matthieu Geist, Fabrice Lefevre, and Olivier Pietquin. 2011 · 2011
Cited alongside, same era.
Toward Learning and Evaluation of Dialogue Policies with Text Examples
David DeVault, Anton Leuski, and Kenji Sagae. 2011 · 2011
Cited alongside, same era.
Zachary C. Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng. 2017 · 2017
Later among the works it cites.
Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models
Bing Liu and Ian Lane. 2017 · 2017
Later among the works it cites.
A Laplacian Framework for Option Discovery in Reinforcement Learning
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling. 2017 · 2017
Later among the works it cites.
Neural Belief Tracker: Data-Driven Dialogue State Tracking
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017 · 2017
Later among the works it cites.
Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Later among the works it cites.
Sample-efficient Actor-Critic Reinforcement Learning with Supervised Data for Dialogue Management
Pei-Hao Su, Pawel Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young. 2017 · 2017
Later among the works it cites.
Pydial: A multi-domain statistical dialogue system toolkit
Stefan Ultes, Lina M Rojas Barahona, Pei-Hao Su, David Vandyke, Dongho Kim, Inigo Casanueva, Paweł Budzianowski, Nikola Mrkšić, Tsung-Hsien Wen, Milica Gasic, et al. 2017 · 2017
Later among the works it cites.
Sequence modeling via segmentations
Chong Wang, Yining Wang, Po-Sen Huang, Abdelrahman Mohamed, Dengyong Zhou, and Li Deng. 2017 · 2017
Later among the works it cites.
Policy Adaptation for Deep Reinforcement Learning-Based Dialogue Management
Lu Chen, Cheng Chang, Zhi Chen, Bowen Tan, Milica Gašić, and Kai Yu. 2018 · 2018
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2018 · 2018
Later among the works it cites.
Autonomous sub-domain modeling for dialogue policy with hierarchical deep reinforcement learning
Giovanni Yoko Kristianto, Huiwen Zhang, Bin Tong, Makoto Iwayama, and Yoshiyuki Kobayashi. 2018 · 2018
Later among the works it cites.
Adversarial Learning of Task-Oriented Neural Dialog Models
Bing Liu and Ian Lane. 2018 · 2018
Later among the works it cites.
Goal-oriented Dialogue Policy Learning from Failures
Keting Lu, Shiqi Zhang, and Xiaoping Chen. 2018 · 2018
Later among the works it cites.
Discriminative Deep Dyna-Q: Robust Planning for Dialogue Policy Learning
Shang-Yu Su, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yun-Nung Chen. 2018 · 2018
Later among the works it cites.
Subgoal Discovery for Hierarchical Dialogue Policy Learning
Da Tang, Xiujun Li, Jianfeng Gao, Chong Wang, Lihong Li, and Tony Jebara. 2018 · 2018
Later among the works it cites.
Sample Efficient Deep Reinforcement Learning for Dialogue Systems with Large Action Spaces
Gellért Weisz, Paweł Budzianowski, Pei-Hao Su, and Milica Gašić. 2018 · 2018
Later among the works it cites.
Yuexin Wu, Xiujun Li, Jingjing Liu, Jianfeng Gao, and Yiming Yang. 2018 · 2018
Later among the works it cites.
A Survey on Reinforcement Learning for Dialogue Systems
Isabella Graßl. 2019 · 2019
Later among the works it cites.
Collaborative multi-agent dialogue model training via reinforcement learning
Alexandros Papangelis, Yi-Chia Wang, Piero Molino, and Gokhan Tur. 2019 · 2019
Later among the works it cites.
Guided Dialog Policy Learning: Reward Estimation for Multi-Domain Task-Oriented Dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 2019
Later among the works it cites.
Multi-Action Dialog Policy Learning with Interactive Human Teaching
Megha Jhunjhunwala, Caleb Bryant, and Pararth Shah. 2020 · 2020
Later among the works it cites.
Towards Sentiment-Aware Multi-Modal Dialogue Policy Learning
Tulika Saha, Sriparna Saha, and Pushpak Bhattacharyya. 2020 · 2020
Later among the works it cites.
Goal-Oriented Chatbot Dialog Management Bootstrapping with Transfer Learning
Vladimir Ilievski, Claudiu Musat, Andreea Hossmann, and Michael Baeriswyl. 2018 · 2021
Later among the works it cites.
Learning dialogue strategies within the Markov decision process framework
E. Levin, R. Pieraccini, and W. Eckert. 1997 · 2021
Later among the works it cites.
Cross-domain Dialogue Policy Transfer via Simultaneous Speech-act and Slot Alignment
Kaixiang Mo, Yu Zhang, Qiang Yang, and Pascale Fung. 2018 · 2021
Later among the works it cites.
Learning agents for uncertain environments (extended abstract)
Stuart Russell. 1998 · 2021
Later among the works it cites.
Pei-Hao Su, David Vandyke, Milica Gasic, Dongho Kim, Nikola Mrksic, Tsung-Hsien Wen, and Steve Young. 2015a · 2021
Later among the works it cites.
A collaborative multi-agent reinforcement learning framework for dialog action decomposition
Huimin Wang and Kam-Fai Wong. 2021 · 2021
Later among the works it cites.