Fetching the paper…
Reading the bibliography…
When learning task-oriented dialogue (ToD) agents, reinforcement learning (RL) techniques can naturally be utilized to train dialogue strategies to achieve user-specific goals.
Alternating recurrent dialog model with large-scale pre-trained language models
Qingyang Wu, Yichi Zhang, Yu Li, and Zhou Yu · 1910
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Spoken natural language dialog systems: A practical approach
Ronnie W Smith and D Richard Hipp · 1994
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Stuart Russell · 1998
Earlier work this paper cites.
Support vector learning for ordinal regression
Ralf Herbrich, Thore Graepel, and Klaus Obermayer · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
An efficient boosting algorithm for combining preferences
Yoav Freund, Raj Iyer, Robert E Schapire, and Yoram Singer · 2003
Earlier work this paper cites.
Learning to rank using gradient descent
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender · 2005
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li · 2007
Earlier work this paper cites.
Partially observable markov decision processes for spoken dialog systems
Jason D Williams and Steve Young · 2007
Earlier work this paper cites.
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li · 2008
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu · 2009
Earlier work this paper cites.
Reinforcement learning of argumentation dialogue policies in negotiation
Kallirroi Georgila and David Traum · 2011
Earlier work this paper cites.
Spoken language understanding: Systems for extracting semantic information from speech
Gokhan Tur and Renato De Mori · 2011
Earlier work this paper cites.
Batch Reinforcement Learning , pp. 45–73
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 2012
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-Hao Su, David Vandyke, and Steve Young · 2015
Earlier work this paper cites.
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Earlier work this paper cites.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrksic, Milica Gasic, Lina M Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
A survey on dialogue systems: Recent advances and new frontiers
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong · 2017
Cited alongside, same era.
The importance of pessimism in fixed-dataset policy optimization
Jacob Buckman, Carles Gelada, and Marc G Bellemare · 2020
Later among the works it cites.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić · 2020
Later among the works it cites.
Bayesian attention modules
Xinjie Fan, Shujian Zhang, Bo Chen, and Mingyuan Zhou · 2020
Later among the works it cites.
End-to-end neural pipeline for goal-oriented dialogue systems using gpt-2
Donghoon Ham, Jeong-Gwan Lee, Youngsoo Jang, and Kee-Eung Kim · 2020
Later among the works it cites.
A simple language model for task-oriented dialogue
Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher · 2020
Later among the works it cites.
Human-centric dialog training via offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li · 2018
Cited alongside, same era.
Playing 20 question game with policy-based reinforcement learning
Huang Hu, Xianchao Wu, Bingfeng Luo, Chongyang Tao, Can Xu, Wei Wu, and Zhan Chen · 2018
Cited alongside, same era.
Curriculum learning based on reward sparseness for deep reinforcement learning of task completion dialogue management
Atsushi Saito · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Shane Gu, and Rosalind Picard · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Guided dialog policy learning without adversarial learning in the loop
Ziming Li, Sungjin Lee, Baolin Peng, Jinchao Li, Julia Kiseleva, Maarten de Rijke, Shahin Shayandeh, and Jianfeng Gao · 2020
Later among the works it cites.
MinTL: Minimalist transfer learning for task-oriented dialogue systems
Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, and Pascale Fung · 2020
Later among the works it cites.
Escaping the gravitational pull of softmax
Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans · 2020
Later among the works it cites.
Task-oriented dialog systems that consider multiple appropriate responses under the same context
Yichi Zhang, Zhijian Ou, and Zhou Yu · 2020
Later among the works it cites.
Convlab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems
Qi Zhu, Zheng Zhang, Yan Fang, Xiang Li, Ryuichi Takanobu, Jinchao Li, Baolin Peng, Jianfeng Gao, Xiaoyan Zhu, and Minlie Huang · 2020
Later among the works it cites.
Adversarial intrinsic motivation for reinforcement learning
Ishan Durugkar, Mauricio Tec, Scott Niekum, and Peter Stone · 2021
Later among the works it cites.
Contextual dropout: An efficient sample-dependent dropout module
Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, and Mingyuan Zhou · 2021
Later among the works it cites.
A Minimalist Approach to Offline Reinforcement Learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Soloist: Buildingtask bots at scale with transfer learning and machine teaching
Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao · 2021
Later among the works it cites.
Annotation inconsistency and entity bias in multiwoz
Kun Qian, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu, and Chinnadhurai Sankar · 2021
Later among the works it cites.
Causal-aware safe policy improvement for task-oriented dialogue
Govardana Sachithanandam Ramachandran, Kazuma Hashimoto, and Caiming Xiong · 2021
Later among the works it cites.
Ubar: Towards fully end-to-end task-oriented dialog system with gpt-2
Yunyi Yang, Yunhao Li, and Xiaojun Quan · 2021
Later among the works it cites.
Galaxy: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, et al · 2022
Later among the works it cites.
GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim · 2022
Later among the works it cites.
Wai-Chung Kwan, Hongru Wang, Huimin Wang, and Kam-Fai Wong · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning
Siddharth Verma, Justin Fu, Mengjiao Yang, and Sergey Levine · 2022
Later among the works it cites.