Fetching the paper…
Reading the bibliography…
Dialogue policy optimization often obtains feedback until task completion in task-oriented dialogue systems.
Stochastic prediction of multi-agent interactions from partial observations
Chen Sun, Per Karlsson, Jiajun Wu, Joshua B Tenenbaum, and Kevin Murphy. 2019 · 1902
Earlier work this paper cites.
Unsupervised learning of object structure and dynamics from videos
Matthias Minderer, Chen Sun, Ruben Villegas, Forrester Cole, Kevin Murphy, and Honglak Lee. 2019 · 1906
Earlier work this paper cites.
Budgeted policy learning for task-oriented dialogue systems
Zhirui Zhang, Xiujun Li, Jianfeng Gao, and Enhong Chen. 2019 · 1906
Earlier work this paper cites.
Mala: Cross-domain dialogue generation with action learning
Xinting Huang, Jianzhong Qi, Yu Sun, and Rui Zhang. 2019a · 1912
Earlier work this paper cites.
Semi-supervised stochastic multi-domain learning using variational inference
Yitong Li, Timothy Baldwin, and Trevor Cohn. 2019a · 1934
Earlier work this paper cites.
On-line policy optimisation of bayesian spoken dialogue systems via human interaction
Milica Gašić, Catherine Breslin, Matthew Henderson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young. 2013 · 2013
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee. 2013 · 2013
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Durk P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling. 2014 · 2014
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Distributional smoothing with virtual adversarial training
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Ken Nakae, and Shin Ishii. 2015 · 2015
Earlier work this paper cites.
Reward shaping with recurrent neural networks for speeding up on-line policy learning in spoken dialogue systems
Pei-Hao Su, David Vandyke, Milica Gasic, Nikola Mrksic, Tsung-Hsien Wen, and Steve Young. 2015 · 2015
Earlier work this paper cites.
On-line active reward learning for policy optimisation in spoken dialogue systems
Pei-Hao Su, Milica Gasic, Nikola Mrkšić, Lina M Rojas Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
Contextual intent tracking for personal assistants
Yu Sun, Nicholas Jing Yuan, Yingzi Wang, Xing Xie, Kieran McDonald, and Rui Zhang. 2016 · 2016
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
End-to-end task-completion neural dialogue systems
Xuijun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018 · 2018
Later among the works it cites.
Deconvolutional latent-variable model for text sequence matching
Dinghan Shen, Yizhe Zhang, Ricardo Henao, Qinliang Su, and Lawrence Carin. 2018 · 2018
Later among the works it cites.
Kdgan: knowledge distillation with generative adversarial networks
Xiaojie Wang, Rui Zhang, Yu Sun, and Jianzhong Qi. 2018 · 2018
Later among the works it cites.
Unsupervised text style transfer using language models as discriminators
Zichao Yang, Zhiting Hu, Chris Dyer, Eric P Xing, and Taylor Berg-Kirkpatrick. 2018 · 2018
Later among the works it cites.
Learning action representations for reinforcement learning
Yash Chandak, Georgios Theocharous, James Kostas, Scott Jordan, and Philip Thomas. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Collaborative intent prediction with real-time contextual data
Yu Sun, Nicholas Jing Yuan, Xing Xie, Kieran McDonald, and Rui Zhang. 2017 · 2017
Cited alongside, same era.
Reward-balancing for statistical spoken dialogue systems using multi-objective reinforcement learning
Stefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Nikola Mrkšić, Lina M Rojas Barahona, Pei-Hao Su, Tsung-Hsien Wen, Milica Gasic, and Steve Young. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Variational autoencoder for semi-supervised text classification
Weidi Xu, Haoze Sun, Chao Deng, and Ying Tan. 2017 · 2017
Cited alongside, same era.
Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018 · 2018
Cited alongside, same era.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2018 · 2018
Cited alongside, same era.
Adversarial learning of task-oriented neural dialog models
Bing Liu and Ian Lane. 2018 · 2018
Cited alongside, same era.
Runyu Chen, Qili Wang, and Wen I Xu. 2019 · 2019
Later among the works it cites.
A cross-sentence latent variable model for semi-supervised text sequence matching
Jihun Choi, Taeuk Kim, and Sang-goo Lee. 2019 · 2019
Later among the works it cites.
Label propagation for deep semi-supervised learning
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondřej Chum. 2019 · 2019
Later among the works it cites.
Convlab: Multi-domain end-to-end dialog system platform
Sungjin Lee, Qi Zhu, Ryuichi Takanobu, Xiang Li, Yaoqin Zhang, Zheng Zhang, Jinchao Li, Baolin Peng, Xiujun Li, Minlie Huang, and Jianfeng Gao. 2019 · 2019
Later among the works it cites.
Adversarial sampling and training for semi-supervised information retrieval
Dae Hoon Park and Yi Chang. 2019 · 2019
Later among the works it cites.
Guided dialog policy learning: Reward estimation for multi-domain task-oriented dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 2019
Later among the works it cites.
Rethinking action spaces for reinforcement learning in end-to-end dialog agents with latent variable models
Tiancheng Zhao, Kaige Xie, and Maxine Eskenazi. 2019 · 2019
Later among the works it cites.