Fetching the paper…
Reading the bibliography…
Despite recent advancements in language models (LMs), their application to dialogue management (DM) problems and ability to carry on rich conversations remain a challenge.
A maximum likelihood approach to continuous speech recognition
L. Bahl, F. Jelinek, and R. Mercer · 1983
Earlier work this paper cites.
A stochastic model of computer-human interaction for learning dialogue strategies
E. Levin and R. Pieraccini · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email
M. Walker · 2000
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
S. Singh, D. Litman, M. Kearns, and M. Walker · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
E. Greensmith, P. Bartlett, and J. Baxter · 2004
Earlier work this paper cites.
Partially observable markov decision processes for spoken dialog systems
J. Williams and S. Young · 2007
Earlier work this paper cites.
Comparing measures of sparsity
N. Hurley and S. Rickard · 2009
Earlier work this paper cites.
The hidden information state model: A practical framework for pomdp-based spoken dialogue management
S. Young, M. Gašić, S. Keizer, F. Mairesse, J. Schatzmann, B. Thomson, and K. Yu · 2010
Earlier work this paper cites.
C. Danescu-Niculescu-Mizil and L. Lee · 2011
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
M. Gašić, F. Jurčíček, B. Thomson, K. Yu, and S. Young · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. Le · 2014
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan · 2015
Earlier work this paper cites.
Sample-efficient deep reinforcement learning for dialog control
K. Asadi and J. Williams · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. Kulkarni D, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Earlier work this paper cites.
Policy networks with two-stage training for dialogue systems
M. Fatemi, L. Asri, H. Schulz, J. He, and K. Suleman · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
J. Li, W. Monroe, A. Ritter, M. Galley, J. Gao, and D. Jurafsky · 2016
Earlier work this paper cites.
Building end-to-end dialogue systems using generative hierarchical neural network models
I. Serban, A. Sordoni, Y. Bengio, A. Courville, and J. Pineau · 2016
Earlier work this paper cites.
Controlling linguistic style aspects in neural language generation
J. Ficler and Y. Goldberg · 2017
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
J. Gu, K. Cho, and V. Li · 2017
Earlier work this paper cites.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, S. Jean, A. Ritter, and D. Jurafsky · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
J. Schulman, X. Chen, and P. Abbeel · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Understanding disentangling in β \beta -vae
C. Burgess, I. Higgins, A. Pal, L. Matthey, N. Watters, G. Desjardins, and A. Lerchner · 2018
Cited alongside, same era.
D. Cer, Y. Yang, S. Kong, N. Hua, N. Limtiaco, R. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al · 2018
Generating diverse high-resolution images with vq-vae
A. Razavi, A. van den Oord, and O. Vinyals · 2019
Later among the works it cites.
Do neural dialog systems use the conversation history effectively? an empirical study
C. Sankar, S. Subramanian, C. Pal, S. Chandar, and Y. Bengio · 2019
Later among the works it cites.
Can unconditional language models recover arbitrary sentences?
N. Subramani, S. Bowman, and K. Cho · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Later among the works it cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Y. Zhang, S. Sun, M. Galley, Y. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A stable and effective learning strategy for trainable greedy decoding
Y. Chen, V. Li, K. Cho, and S. Bowman · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. Van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Learning to write with cooperative discriminators
A. Holtzman, J. Buys, M. Forbes, A. Bosselut, D. Golub, and Y. Choi · 2018
Cited alongside, same era.
A hierarchical latent structure for variational conversation modeling
Y. Park, J. Cho, and G. Kim · 2018
Cited alongside, same era.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
P. Shah, D. Hakkani-Tur, B. Liu, and G. Tür · 2018
Cited alongside, same era.
Later among the works it cites.
T. Zhao, K. Xie, and M. Eskenazi · 2019
Later among the works it cites.
Fine-tuning language models from human preferences
D. Ziegler, N. Stiennon, J. Wu, T. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2020
Later among the works it cites.
GPT-3: Its nature, scope, limits, and consequences
L. Floridi and M. Chiriatti · 2020
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Few-shot natural language generation for task-oriented dialog
B. Peng, C. Zhu, C. Li, X. Li, J. Li, M. Zeng, and J. Gao · 2020
Later among the works it cites.
Hierarchical reinforcement learning for open-domain dialog
A. Saleh, N. Jaques, A. Ghandeharioun, J. Shen, and R. Picard · 2020
Later among the works it cites.
Generating empathetic responses by looking ahead the user’s sentiment
J. Shin, P. Xu, A. Madotto, and P. Fung · 2020
Later among the works it cites.
Predictive coding for locally-linear control
R. Shu, T. Nguyen, Y. Chow, T. Pham, K. Than, M. Ghavamzadeh, S. Ermon, and H. Bui · 2020
Later among the works it cites.
The design and implementation of xiaoice, an empathetic social chatbot
L. Zhou, J. Gao, D. Li, and H. Shum · 2020
Later among the works it cites.
How bpe affects memorization in transformers
E. Kharitonov, M. Baroni, and D. Hupkes · 2021
Later among the works it cites.
An improved aspect-category sentiment analysis model for text sentiment analysis based on roberta
W. Liao, B. Zeng, X. Yin, and P. Wei · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
M. Yang and O. Nachum · 2021
Later among the works it cites.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning
S. Verma, J. Fu, M. Yang, and S. Levine · 2022
Closest in time.