Fetching the paper…
Reading the bibliography…
Goal-oriented dialogue systems face a trade-off between fluent language generation and task-specific control.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum. 2019b · 1911
Earlier work this paper cites.
Recommendation as a communication game: Self-supervised bot-play for goal-oriented dialogue
Dongyeop Kang, Anusha Balakrishnan, Pararth Shah, Paul A Crook, Y-Lan Boureau, and Jason Weston. 2019 · 1961
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau. 1989 · 1989
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling. 1993 · 1993
Earlier work this paper cites.
Spoken natural language dialog systems: A practical approach
Ronnie W Smith and D Richard Hipp. 1994 · 1994
Earlier work this paper cites.
User modeling for spoken dialogue system evaluation
Wieland Eckert, Esther Levin, and Roberto Pieraccini. 1997 · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra. 1998 · 1998
Earlier work this paper cites.
Reinforcement learning for spoken dialogue systems
Satinder Singh, Michael Kearns, Diane Litman, and Marilyn Walker. 1999 · 1999
Earlier work this paper cites.
A stochastic model of human-machine interaction for learning dialog strategies
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000 · 2000
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Rewriting history with inverse rl: Hindsight inference for policy improvement
Benjamin Eysenbach, Xinyang Geng, Sergey Levine, and Ruslan Salakhutdinov. 2020 · 2002
Earlier work this paper cites.
Spoken dialogue technology: enabling the conversational user interface
Michael F McTear. 2002 · 2002
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker. 2002 · 2002
Earlier work this paper cites.
Developing a flexible spoken dialog system using simulation
Grace Chung. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2005
Earlier work this paper cites.
Soloist: Few-shot task-oriented dialog with a single pre-trained auto-regressive model
Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2020 · 2005
Earlier work this paper cites.
User simulation for spoken dialogue systems: Learning and evaluation
Kallirroi Georgila, James Henderson, and Oliver Lemon. 2006 · 2006
Earlier work this paper cites.
Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu. 2020a · 2006
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Cited alongside, same era.
A human-computer dialogue system for educational debate: A computational dialectics approach
Tangming Yuan, David Moore, and Alec Grierson. 2008 · 2008
Cited alongside, same era.
Representing the reinforcement learning state in a negotiation dialogue
Peter A Heeman. 2009 · 2009
Cited alongside, same era.
Dialogue-oriented review summary generation for spoken dialogue recommendation systems
Jingjing Liu, Stephanie Seneff, and Victor Zue. 2010 · 2010
Cited alongside, same era.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurčíček, Blaise Thomson, Kai Yu, and Steve Young. 2011 · 2011
Cited alongside, same era.
Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018 · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kallirroi Georgila and David Traum. 2011 · 2011
Cited alongside, same era.
Sample-efficient batch reinforcement learning for dialogue management optimization
Olivier Pietquin, Matthieu Geist, Senthilkumar Chandramohan, and Hervé Frezza-Buet. 2011 · 2011
Cited alongside, same era.
Lstm neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney. 2012 · 2012
Cited alongside, same era.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Cited alongside, same era.
The dialog state tracking challenge series
Jason D Williams, Matthew Henderson, Antoine Raux, Blaise Thomson, Alan Black, and Deepak Ramachandran. 2014 · 2014
Cited alongside, same era.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Layla El Asri, Jing He, and Kaheer Suleman. 2016 · 2016
Cited alongside, same era.
Sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings
Ondřej Dušek and Filip Jurcicek. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun. 2018 · 2018
Later among the works it cites.
Airdialogue: An environment for goal-oriented dialogue research
Wei Wei, Quoc Le, Andrew Dai, and Jia Li. 2018 · 2018
Later among the works it cites.
Semantically conditioned dialog response generation via hierarchical disentangled self-attention
Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, and William Yang Wang. 2019 · 2019
Later among the works it cites.
Learning to reach goals without reinforcement learning
Dibya Ghosh, Abhishek Gupta, Justin Fu, Ashwin Reddy, Coline Devin, Benjamin Eysenbach, and Sergey Levine. 2019 · 2019
Later among the works it cites.
Tiancheng Zhao, Kaige Xie, and Maxine Eskenazi. 2019 · 2019
Later among the works it cites.
Airconcierge: Generating task-oriented dialogue via efficient large-scale knowledge retrieval
Chieh-Yang Chen, Pei-Hsin Wang, Shih-Chieh Chang, Da-Cheng Juan, Wei Wei, and Jia-Yu Pan. 2020 · 2020
Later among the works it cites.
A simple language model for task-oriented dialogue
Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020 · 2020
Later among the works it cites.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet. 2020 · 2020
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Task-oriented dialog systems that consider multiple appropriate responses under the same context
Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020 · 2020
Later among the works it cites.
Towards automatic evaluation of dialog systems: A model-free off-policy evaluation approach
Haoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao, and Wei Wei. 2021 · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2021 · 2021
Later among the works it cites.
Tapex: Table pre-training via learning a neural sql executor
Qian Liu, Bei Chen, Jiaqi Guo, Zeqi Lin, and Jian-guang Lou. 2021 · 2021
Later among the works it cites.
When should we prefer offline reinforcement learning over behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine. 2022 · 2022
Closest in time.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018 · 2069
Closest in time.