Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) is poised to revolutionise the field of artificial intelligence (AI) by endowing autonomous systems with high levels of understanding of the real world.
B. F. Skinner, “Verbal behavior. new york: appleton-century-crofts,” Richard-Amato, P.(1996) , vol. 11, 1957
1957
Earlier work this paper cites.
R. Bellman, “Dynamic programming,” Science , vol. 153, no. 3731, 1966
1966
Earlier work this paper cites.
M. J. Steedman, “A generative grammar for jazz chord sequences,” Music Perception: An Interdisciplinary Journal , vol. 2, no. 1, 1984
1984
Earlier work this paper cites.
K. Ebcioğlu, “An expert system for harmonizing four-part chorales,” Computer Music Journal , vol. 12, no. 3, 1988
1988
Earlier work this paper cites.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural computation , vol. 1, no. 4, 1989
1989
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based CSR corpus,” in Workshop on Speech and Natural Language . ACL, 1992
1992
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 1, 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3-4, 1992
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,” NASA STI/Recon technical report n , vol. 93, 1993
1993
Earlier work this paper cites.
M. L. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in Machine learning proceedings 1994 . Elsevier, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, 1997
1997
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Transactions Signal Process. , vol. 45, no. 11, 1997
1997
Earlier work this paper cites.
R. S. Sutton, A. G. Barto et al. , Introduction to reinforcement learning . MIT press Cambridge, 1998, vol. 135
1998
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-Critic agorithms,” in Neural Information Processing Systems (NIPS) , 1999
1999
Earlier work this paper cites.
V. W. Zue and J. R. Glass, “Conversational interfaces: advances and challenges,” IEEE , vol. 88, no. 8, 2000
2000
Earlier work this paper cites.
E. Levin, R. Pieraccini, and W. Eckert, “A stochastic model of human-machine interaction for learning dialog strategies,” IEEE Transactions Speech Audio Process. , vol. 8, no. 1, 2000
2000
Earlier work this paper cites.
S. P. Singh, M. J. Kearns, D. J. Litman, and M. A. Walker, “Reinforcement learning for spoken dialogue systems,” in Advances in Neural Information Processing Systems (NIPS) , 2000
2000
Earlier work this paper cites.
I.-T. Recommendation, “Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Rec. ITU-T P. 862 , 2001
2001
Earlier work this paper cites.
S. Singh, D. Litman, M. Kearns, and M. Walker, “Optimizing dialogue management with reinforcement learning: Experiments with the njfun system,” Journal of Artificial Intelligence Research , vol. 16, 2002
2002
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in IEEE International Conference on Robotics and Automation (ICRA) , vol. 3, 2004
2004
Earlier work this paper cites.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, and B. Weiss, “A database of german emotional speech,” in European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
M. Allan and C. Williams, “Harmonising chorales by probabilistic inference,” in Advances in Neural Information Processing Systems (NIPS) , 2005
2005
Earlier work this paper cites.
A. Y. Ng, A. Coates, M. Diel, V. Ganapathi, J. Schulte, B. Tse, E. Berger, and E. Liang, “Autonomous inverted helicopter flight via reinforcement learning,” in Experimental robotics IX . Springer, 2006
2006
Earlier work this paper cites.
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman, “Pac model-free reinforcement learning,” in International Conference on Machine Learning (ICML) , 2006
2006
Earlier work this paper cites.
J. Schatzmann, K. Weilhammer, M. N. Stuttle, and S. J. Young, “A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies,” Knowledge Eng. Review , vol. 21, no. 2, 2006
2006
Earlier work this paper cites.
T. Paek, “Reinforcement learning for spoken dialogue systems: Comparing strengths and weaknesses for practical deployment,” in Proc. Dialog-on-Dialog Workshop, Interspeech . Citeseer, 2006
2006
Earlier work this paper cites.
M. A. Goodrich and A. C. Schultz, “Human-robot interaction: a survey,” Foundations and trends in human-computer interaction , vol. 1, no. 3, 2007
2007
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, 2008
2008
Earlier work this paper cites.
A.-r. Mohamed, G. Dahl, and G. Hinton, “Deep belief networks for phone recognition,” in NIPS workshop on deep learning for speech recognition and related applications , 2009
2009
Earlier work this paper cites.
H. Cuayáhuitl, “Hierarchical reinforcement learning for spoken dialogue systems,” Ph.D. dissertation, University of Edinburgh, 2009
2009
Earlier work this paper cites.
H. Cuayáhuitl, S. Renals, O. Lemon, and H. Shimodaira, “Evaluation of a hierarchical reinforcement learning spoken dialogue system,” Comput. Speech Lang. , vol. 24, no. 2, 2010
2010
Earlier work this paper cites.
G. McKeown, M. Valstar, R. Cowie, M. Pantic, and M. Schroder, “The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,” IEEE transactions on affective computing , vol. 3, no. 1, 2011
2011
Earlier work this paper cites.
V. Emiya, E. Vincent, N. Harlander, and V. Hohmann, “Subjective and objective quality assessment of audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 7, 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, 2012
2012
Earlier work this paper cites.
S. Lange, M. A. Riedmiller, and A. Voigtländer, “Autonomous reinforcement learning on raw visual input data in a real world application,” in International Joint Conference on Neural Networks (IJCNN), Brisbane, Australia, June 10-15, 2012 . IEEE, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS) , 2012
2012
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” Workshop on Representation Learning, International Conference of Machine Learning (ICML) 2012 , 2012
2012
Earlier work this paper cites.
M. Shannon, H. Zen, and W. Byrne, “Autoregressive models for statistical parametric speech synthesis,” IEEE transactions on audio, speech, and language processing , vol. 21, no. 3, 2012
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “TED-LIUM: an automatic speech recognition dedicated corpus.” in LREC , 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. P. Rath, D. Povey, K. Veselý, and J. Cernocký, “Improved feature processing for deep neural networks,” in Interspeech . ISCA, 2013
2013
Earlier work this paper cites.
D. Ameixa, L. Coheur, and R. A. Redol, “From subtitles to human interactions: introducing the subtle corpus,” Tech. rep., INESC-ID (November 2014), Tech. Rep., 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The diverse environments multi-channel acoustic noise database: A database of multichannel environmental noise recordings,” The Journal of the Acoustical Society of America , vol. 133, no. 5, 2013
2013
Earlier work this paper cites.
M. Gašić and S. Young, “Gaussian processes for POMDP-based dialogue manager optimization,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 1, 2013
2013
Earlier work this paper cites.
B. Li, Y. Tsao, and K. C. Sim, “An investigation of spectral restoration algorithms for deep neural networks based noise robust speech recognition.” in Interspeech , 2013
2013
Earlier work this paper cites.
N. Howard and E. Cambria, “Intention awareness: Improving upon situation awareness in human-centric environments,” Human-centric Computing and Information Sciences , vol. 3, no. 9, 2013
2013
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The Arcade learning environment: An evaluation platform for general agents,” J. Artif. Intell. Res. , vol. 47, 2013
2013
Earlier work this paper cites.
J. Schlüter and S. Böck, “Improved musical onset detection with convolutional neural networks,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014
2014
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on audio, speech, and language processing , vol. 22, no. 10, 2014
2014
Earlier work this paper cites.
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems (NIPS) , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NIPS) , 2014
2014
Earlier work this paper cites.
M. Henderson, B. Thomson, and J. D. Williams, “The second dialog state tracking challenge,” in Proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (SIGDIAL) , 2014, pp. 263–272
2014
Earlier work this paper cites.
——, “The third dialog state tracking challenge,” in 2014 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2014, pp. 324–329
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Li, A. Mohamed, G. Zweig, and Y. Gong, “LSTM time and frequency recurrence for automatic speech recognition,” in IEEE workshop on automatic speech recognition and understanding (ASRU) , 2015
2015
Earlier work this paper cites.
M. Hausknecht and P. Stone, “Deep recurrent Q-learning for partially observable MDPs,” in AAAI Fall Symposium Series , 2015
2015
Earlier work this paper cites.
I. Sorokin, A. Seleznev, M. Pavlov, A. Fedorov, and A. Ignateva, “Deep attention recurrent Q-network,” Deep Reinforcement Learning Workshop, NIPS , 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning (ICML) , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015
2015
Earlier work this paper cites.
N. Dethlefs and H. Cuayáhuitl, “Hierarchical reinforcement learning for situated natural language generation,” Nat. Lang. Eng. , vol. 21, no. 3, 2015
2015
Earlier work this paper cites.
J. Li, L. Deng, R. Haeb-Umbach, and Y. Gong, Robust automatic speech recognition: a bridge to practical applications . Academic Press, 2015
2015
Earlier work this paper cites.
D. Baby, J. F. Gemmeke, T. Virtanen et al. , “Exemplar-based speech enhancement for deep neural network based automatic speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, 2016
2016
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. N. Sainath and B. Li, “Modeling time-frequency patterns with lstm vs. convolutional architectures for lvcsr tasks,” in Interspeech , 2016
2016
Earlier work this paper cites.
Y. Qian, M. Bi, T. Tan, and K. Yu, “Very deep convolutional neural networks for noise robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 12, 2016
2016
Earlier work this paper cites.
L. Lu, X. Zhang, and S. Renals, “On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Earlier work this paper cites.
N. Jaitly, Q. V. Le, O. Vinyals, I. Sutskever, D. Sussillo, and S. Bengio, “An online sequence-to-sequence model using partial conditioning,” in Advances in Neural Information Processing Systems (NIPS) , 2016
2016
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in AAAI Conference , 2016
2016
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” International Conference on Learning Representations (ICLR) , 2016
2016
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning (ICML) , 2016
2016
Earlier work this paper cites.
J. Oh, V. Chockalingam, H. Lee et al. , “Control of memory, active perception, and action in minecraft,” in International Conference on Machine Learning , 2016
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning (ICML) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare, “Safe and efficient off-policy reinforcement learning,” in Advances in Neural Information Processing Systems (NIPS) , 2016
2016
Earlier work this paper cites.
A. Vezhnevets, V. Mnih, S. Osindero, A. Graves, O. Vinyals, J. Agapiou et al. , “Strategic attentive writer for learning macro-actions,” in Advances in Neural Information Processing Systems (NIPS) , 2016
2016
Earlier work this paper cites.
J. D. Williams, A. Raux, and M. Henderson, “The dialog state tracking challenge series: A review,” Dialogue Discourse , vol. 7, no. 3, 2016
2016
Earlier work this paper cites.
C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, and E. M. Provost, “MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception,” IEEE Transactions on Affective Computing , vol. 8, no. 1, 2016
2016
Cited alongside, same era.
B. Krueger, “Classical piano midi page,” 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. K. Misra, J. Sung, K. Lee, and A. Saxena, “Tell me dave: Context-sensitive grounding of natural language to manipulation instructions,” Int. J. Robotics Res. , vol. 35, no. 1-3, 2016
2016
Cited alongside, same era.
D. Lawson, C.-C. Chiu, G. Tucker, C. Raffel, K. Swersky, and N. Jaitly, “Learning hard alignments with variational inference,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Later among the works it cites.
G. Weisz, P. Budzianowski, P.-H. Su, and M. Gašić, “Sample efficient deep reinforcement learning for dialogue systems with large action spaces,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 11, 2018
2018
Later among the works it cites.
J. Zhang, T. Zhao, and Z. Yu, “Multimodal hierarchical reinforcement learning policy for task-oriented visual dialog,” in Annual SIGdial Meeting on Discourse and Dialogue, Melbourne, Australia, July 12-14, 2018 , K. Komatani, D. J. Litman, K. Yu, L. Cavedon, M. Nakano, and A. Papangelis, Eds. ACL, 2018
2018
Later among the works it cites.
N. Carrara, R. Laroche, J.-L. Bouraoui, T. Urvoy, and O. Pietquin, “Safe transfer learning for dialogue applications,” 2018
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
T. Zhao and M. Eskenazi, “Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning,” in Annual Meeting of the Special Interest Group on Discourse and Dialogue , 2016
2016
Cited alongside, same era.
H. Cuayáhuitl, S. Yu, A. Williamson, and J. Carse, “Deep reinforcement learning for multi-domain dialogue systems,” NIPS Workshop on Deep Reinforcement Learning , 2016
2016
Cited alongside, same era.
P.-H. Su, M. Gasic, N. Mrkšić, L. M. R. Barahona, S. Ultes, D. Vandyke, T.-H. Wen, and S. Young, “On-line active reward learning for policy optimisation in spoken dialogue systems,” in Annual Meeting of the Association for Computational Linguistics (ACL) , 2016
2016
Cited alongside, same era.
M. Fatemi, L. E. Asri, H. Schulz, J. He, and K. Suleman, “Policy networks with two-stage training for dialogue systems,” in Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL) , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Later among the works it cites.
I. Casanueva, P. Budzianowski, P. Su, S. Ultes, L. M. Rojas-Barahona, B. Tseng, and M. Gasic, “Feudal reinforcement learning for dialogue management in large domains,” in North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) , M. A. Walker, H. Ji, and A. Stent, Eds., 2018
2018
Later among the works it cites.
L. Chen, C. Chang, Z. Chen, B. Tan, M. Gasic, and K. Yu, “Policy adaptation for deep reinforcement learning-based dialogue management,” in IEEE International Conference on Acoustics, Speech and Signal ICASSP , 2018
2018
Later among the works it cites.
Z. C. Lipton, X. Li, J. Gao, L. Li, F. Ahmed, and L. Deng, “Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems,” in AAAI Conference on Artificial Intelligence , S. A. McIlraith and K. Q. Weinberger, Eds., 2018
2018
Later among the works it cites.
K. Narasimhan, R. Barzilay, and T. S. Jaakkola, “Grounding language for transfer in deep reinforcement learning,” J. Artif. Intell. Res. , vol. 63, 2018
2018
Later among the works it cites.
B. Peng, X. Li, J. Gao, J. Liu, Y. Chen, and K. Wong, “Adversarial advantage actor-critic model for task-completion dialogue policy learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
H. Zhou, M. Huang, T. Zhang, X. Zhu, and B. Liu, “Emotional chatting machine: Emotional conversation generation with internal and external memory,” in AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
X. Ouyang, S. Nagisetty, E. G. H. Goh, S. Shen, W. Ding, H. Ming, and D.-Y. Huang, “Audio-visual emotion recognition with capsule-like feature representation and model-based reinforcement learning,” in 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia) . IEEE, 2018, pp. 1–6
2018
Later among the works it cites.
E. Lakomkin, M. A. Zamani, C. Weber, S. Magg, and S. Wermter, “Emorl: continuous acoustic emotion classification using deep reinforcement learning,” in IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Dorfer, F. Henkel, and G. Widmer, “Learning to listen, read, and follow: Score following as a reinforcement learning game,” International Society for Music Information Retrieval Conference , 2018
2018
Later among the works it cites.
H. Yu, H. Zhang, and W. Xu, “Interactive grounded language acquisition and generalization in a 2D world,” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
F. Hill, K. M. Hermann, P. Blunsom, and S. Clark, “Understanding grounded language learning agents,” 2018
2018
Later among the works it cites.
——, “Deep reinforcement learning for audio-visual gaze control,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2018
2018
Later among the works it cites.
M. Clark-Turner and M. Begum, “Deep reinforcement learning of abstract reasoning from demonstrations,” in ACM/IEEE International Conference on Human-Robot Interaction , 2018
2018
Later among the works it cites.
M. Zamani, S. Magg, C. Weber, S. Wermter, and D. Fu, “Deep reinforcement learning using compositional representations for performing instructions,” Paladyn J. Behav. Robotics , vol. 9, no. 1, 2018
2018
Later among the works it cites.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Advances in Neural Information Processing Systems (NIPS) , 2018
2018
Later among the works it cites.
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee, “Sample-efficient reinforcement learning with stochastic ensemble value expansion,” in Advances in Neural Information Processing Systems (NIPS) , 2018
2018
Later among the works it cites.
K. Mo, Y. Zhang, S. Li, J. Li, and Q. Yang, “Personalizing a dialogue system with transfer reinforcement learning,” in AAAI Conference , 2018
2018
Later among the works it cites.
H. Purwins, B. Li, T. Virtanen, J. Schlüter, S.-Y. Chang, and T. Sainath, “Deep learning for audio signal processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 2, 2019
2019
Later among the works it cites.
N. C. Luong, D. T. Hoang, S. Gong, D. Niyato, P. Wang, Y.-C. Liang, and D. I. Kim, “Applications of deep reinforcement learning in communications and networking: A survey,” IEEE Communications Surveys & Tutorials , vol. 21, no. 4, 2019
2019
Later among the works it cites.
N. Mamun, S. Khorram, and J. H. Hansen, “Convolutional neural network-based speech enhancement for cochlear implant recipients,” in Interspeech , 2019
2019
Later among the works it cites.
Y. Chen, Q. Guo, X. Liang, J. Wang, and Y. Qian, “Environmental sound classification with dilated convolutions,” Applied Acoustics , vol. 148, 2019
2019
Later among the works it cites.
R. Liu, J. Yang, and M. Liu, “A new end-to-end long-time speech synthesis system based on tacotron2,” in International Symposium on Signal Processing Systems , 2019
2019
Later among the works it cites.
N. Pham, T. Nguyen, J. Niehues, M. Müller, and A. Waibel, “Very deep self-attention networks for end-to-end speech recognition,” in Interspeech , G. Kubin and Z. Kacic, Eds. ISCA, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter, “Rudder: Return decomposition for delayed rewards,” in Advances in Neural Information Processing Systems (NIPS) , 2019
2019
Later among the works it cites.
B. Ravindran, “Introduction to deep reinforcement learning,” 2019
2019
Later among the works it cites.
Ł. Kaiser, M. Babaeizadeh, P. Miłos, B. Osiński, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine et al. , “Model based reinforcement learning for atari,” in International Conference on Learning Representations , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “MELD: A multimodal multi-party dataset for emotion recognition in conversations,” in Annual Meeting of the Association for Computational Linguistics ACL , 2019
2019
Later among the works it cites.
——, “End-to-end speech recognition sequence training with reinforcement learning,” IEEE Access , vol. 7, 2019
2019
Later among the works it cites.
K. Radzikowski, R. Nowak, L. Wang, and O. Yoshie, “Dual supervised learning for non-native speech recognition,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2019, no. 1, 2019
2019
Later among the works it cites.
Ł. Dudziak, M. S. Abdelfattah, R. Vipperla, S. Laskaridis, and N. D. Lane, “ShrinkML: End-to-end asr model compression using reinforcement learning,” in Interspeech , 2019
2019
Later among the works it cites.
J. Gao, M. Galley, and L. Li, “Neural approaches to conversational AI,” Found. Trends Inf. Retr. , vol. 13, no. 2-3, 2019
2019
Later among the works it cites.
H. Cuayáhuitl, D. Lee, S. Ryu, Y. Cho, S. Choi, S. R. Indurthi, S. Yu, H. Choi, I. Hwang, and J. Kim, “Ensemble-based deep reinforcement learning for chatbots,” Neurocomputing , vol. 366, 2019
2019
Later among the works it cites.
P. Ammanabrolu and M. Riedl, “Transfer in deep reinforcement learning using knowledge graphs,” in Workshop on Graph-Based Methods for Natural Language Processing, TextGraphs@EMNLP , D. Ustalov, S. Somasundaran, P. Jansen, G. Glavas, M. Riedl, M. Surdeanu, and M. Vazirgiannis, Eds. Association for Computational Linguistics, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Sankar and S. Ravi, “Deep reinforcement learning for modeling chit-chat dialog with discrete attributes,” in SIGdial Meeting on Discourse and Dialogue , S. Nakamura, M. Gasic, I. Zuckerman, G. Skantze, M. Nakano, A. Papangelis, S. Ultes, and K. Yoshino, Eds., 2019
2019
Later among the works it cites.
L. Xu, Q. Zhou, K. Gong, X. Liang, J. Tang, and L. Lin, “End-to-end knowledge-routed relational dialogue system for automatic diagnosis,” in AAAI Conference on Artificial Intelligence , 2019
2019
Later among the works it cites.
T. Zhao, K. Xie, and M. Eskénazi, “Rethinking action spaces for reinforcement learning in end-to-end dialog agents with latent variable models,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) , J. Burstein, C. Doran, and T. Solorio, Eds., 2019
2019
Later among the works it cites.
R. Takanobu, H. Zhu, and M. Huang, “Guided dialog policy learning: Reward estimation for multi-domain task-oriented dialog,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 , K. Inui, J. Jiang, V. Ng, and X. Wan, Eds., 2019
2019
Later among the works it cites.
S. Latif, J. Qadir, and M. Bilal, “Unsupervised adversarial domain adaptation for cross-lingual speech emotion recognition,” in International Conference on Affective Computing and Intelligent Interaction (ACII) , 2019
2019
Later among the works it cites.
N. Majumder, S. Poria, D. Hazarika, R. Mihalcea, A. Gelbukh, and E. Cambria, “DialogueRNN: An attentive RNN for emotion detection in conversations,” in AAAI Conference on Artificial Intelligence , vol. 33, 2019
2019
Later among the works it cites.
S. Poria, N. Majumder, R. Mihalcea, and E. Hovy, “Emotion recognition in conversation: Research challenges, datasets, and recent advances,” IEEE Access , vol. 7, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Sangeetha and T. Jayasankar, “Emotion speech recognition based on adaptive fractional deep belief network and reinforcement learning,” in Cognitive Informatics and Soft Computing . Springer, 2019
2019
Later among the works it cites.
Y.-L. Shen, C.-Y. Huang, S.-S. Wang, Y. Tsao, H.-M. Wang, and T.-S. Chi, “Reinforcement learning based speech enhancement for robust speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019
2019
Later among the works it cites.
Q. Lan, J. Tørresen, and A. R. Jensenius, “RaveForce: A deep reinforcement learning environment for music,” in Proc. of the SMC Conferences . Society for Sound and Music Computing, 2019
2019
Later among the works it cites.
F. Henkel, S. Balke, M. Dorfer, and G. Widmer, “Score following as a multi-modal reinforcement learning problem,” Transactions of the International Society for Music Information Retrieval , vol. 2, no. 1, 2019
2019
Later among the works it cites.
A. Sinha, B. Akilesh, M. Sarkar, and B. Krishnamurthy, “Attention based natural language grounding by navigating virtual environment,” in IEEE Winter Conference on Applications of Computer Vision (WACV) , 2019
2019
Later among the works it cites.
S. Lathuilière, B. Massé, P. Mesejo, and R. Horaud, “Neural network based reinforcement learning for audio–visual gaze control in human–robot interaction,” Pattern Recognition Letters , vol. 118, 2019
2019
Later among the works it cites.
N. Hussain, E. Erzin, T. M. Sezgin, and Y. Yemez, “Speech driven backchannel generation using deep q-network for enhancing engagement in human-robot interaction,” in Interspeech , 2019
2019
Later among the works it cites.
——, “Batch recurrent Q-learning for backchannel generation towards engaging agents,” in International Conference on Affective Computing and Intelligent Interaction (ACII) , 2019
2019
Later among the works it cites.
H. Bui and N. Y. Chong, “Autonomous speech volume control for social robots in a noisy environment using deep reinforcement learning,” in IEEE International Conference on Robotics and Biomimetics (ROBIO) , 2019
2019
Later among the works it cites.
P. Hernandez-Leal, B. Kartal, and M. E. Taylor, “A survey and critique of multiagent deep reinforcement learning,” Autonomous Agents and Multi-Agent Systems , vol. 33, no. 6, 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Khan, A. Sohail, U. Zahoora, and A. S. Qureshi, “A survey of the recent architectures of deep convolutional neural networks,” Artif. Intell. Rev. , vol. 53, no. 8, 2020
2020
Later among the works it cites.
S. Latif, J. Qadir, A. Qayyum, M. Usama, and S. Younis, “Speech technology for healthcare: Opportunities, challenges, and state of the art,” IEEE Reviews in Biomedical Engineering , 2020
2020
Later among the works it cites.
S. Latif, “Deep representation learning for improving speech emotion recognition,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Rastogi, X. Zang, S. Sunkara, R. Gupta, and P. Khaitan, “Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset,” in The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI . AAAI Press, 2020
2020
Later among the works it cites.
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux, “WHAMR!: Noisy and reverberant single-channel speech separation,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020
2020
Later among the works it cites.
H. Cuayáhuitl, “A data-efficient deep learning approach for deployable multimodal social robots,” Neurocomputing , vol. 396, 2020
2020
Later among the works it cites.
H. Chung, H.-B. Jeon, and J. G. Park, “Semi-supervised training for sequence-to-sequence speech recognition using reinforcement learning,” in 2020 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2020, pp. 1–6
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
G. Gordon-Hall, P. J. Gorinski, and S. B. Cohen, “Learning dialog policies from weak demonstrations,” in Annual Meeting of the Association for Computational Linguistics ACL , D. Jurafsky, J. Chai, N. Schluter, and J. R. Tetreault, Eds. ACL, 2020
2020
Later among the works it cites.
A. Saleh, N. Jaques, A. Ghandeharioun, J. H. Shen, and R. W. Picard, “Hierarchical reinforcement learning for open-domain dialog,” in AAAI Conference on Artificial Intelligence , 2020
2020
Later among the works it cites.
Z. Wang, S. Ho, and E. Cambria, “A review of emotion sensing: Categorization models and algorithms,” Multimedia Tools and Applications , 2020
2020
Later among the works it cites.
Y. Ma, K. L. Nguyen, F. Xing, and E. Cambria, “A survey on empathetic dialogue systems,” Information Fusion , vol. 64, 2020
2020
Later among the works it cites.
T. Young, V. Pandelea, S. Poria, and E. Cambria, “Dialogue systems with audio context,” Neurocomputing , vol. 388, 2020
2020
Later among the works it cites.
N. Alamdari, E. Lobarinas, and N. Kehtarnavaz, “Personalization of hearing aid compression by human-in-the-loop deep reinforcement learning,” IEEE Access , vol. 8, pp. 203 503–203 515, 2020
2020
Later among the works it cites.
N. Jiang, S. Jin, Z. Duan, and C. Zhang, “Rl-duet: Online music accompaniment generation using deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 01, 2020, pp. 710–718
2020
Later among the works it cites.
S. Gao, W. Hou, T. Tanaka, and T. Shinozaki, “Spoken language acquisition based on reinforcement learning and word unit segmentation,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
T. T. Nguyen, N. D. Nguyen, and S. Nahavandi, “Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,” IEEE transactions on cybernetics , 2020
2020
Later among the works it cites.