Fetching the paper…
Reading the bibliography…
Neural Language Models (NLM), when trained and evaluated with context spanning multiple utterances, have been shown to consistently outperform both conventional n-gram language models and NLMs that use limited context.
R. Kneser and H. Ney, “Improved backing-off for m-gram language modeling.” in ICASSP . IEEE Computer Society, 1995, pp. 181–184
1995
Earlier work this paper cites.
Y. Bengio, R. Ducharme, and P. Vincent, “A neural probabilistic language model,” in Advances in Neural Information Processing Systems , T. Leen, T. Dietterich, and V. Tresp, Eds., vol. 13. MIT Press, 2001
2001
Earlier work this paper cites.
O. Lemon, K. Georgila, J. Henderson, and M. Stuttle, “An ISU dialogue system exhibiting reinforcement learning of dialogue policies: Generic slot-filling in the TALK in-car system,” in Demonstrations , 2006
2006
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur, “Recurrent neural network based language model.” in Interspeech , T. Kobayashi, K. Hirose, and S. Nakamura, Eds., 2010, pp. 1045–1048
2010
Earlier work this paper cites.
P. Aleksic, M. Ghodsi, A. Michaely, C. Allauzen, K. Hall, B. Roark, D. Rybach, and P. Moreno, “Bringing contextual information to google speech recognition,” in Interspeech , 2015
2015
Earlier work this paper cites.
X. Chen, T. Tan, X. Liu, P. Lanchantin, M. Wan, M. J. F. Gales, and P. C. Woodland, “Recurrent neural network language model adaptation for multi-genre broadcast speech recognition,” in Interspeech 2015, Dresden, Germany, September 6-10, 2015 . ISCA, 2015, pp. 3511–3515
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR, San Diego, CA, USA , Y. Bengio and Y. LeCun, Eds., 2015
2015
Earlier work this paper cites.
Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, “Hierarchical attention networks for document classification,” in NAACL-HLT 2016 . San Diego, California: Association for Computational Linguistics, Jun. 2016
2016
Earlier work this paper cites.
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30, 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Gandhe, A. Rastrow, and B. Hoffmeister, “Scalable language model adaptation for spoken dialogue systems,” in SLT Workshop 2018, Athens, Greece . IEEE, 2018, pp. 907–912
2018
Earlier work this paper cites.
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, “The microsoft 2017 conversational speech recognition system,” in ICASSP, Calgary, Canada . IEEE, 2018, pp. 5934–5938
2018
Cited alongside, same era.
W. Xiong, L. Wu, J. Zhang, and A. Stolcke, “Session-level language modeling for conversational speech,” in EMNLP . Brussels, Belgium: ACL, 2018, pp. 2764–2768
2018
Cited alongside, same era.
K. Li, H. Xu, Y. Wang, D. Povey, and S. Khudanpur, “Recurrent neural network language model adaptation for conversational speech recognition,” in Interspeech, Hyderabad, India, 2-6 September 2018 , 2018, pp. 3373–3377
2018
Cited alongside, same era.
A. Raju, B. Hedayatnia, L. Liu, A. Gandhe, C. Khatri, A. Metallinou, A. Venkatesh, and A. Rastrow, “Contextual language model adaptation for conversational agents,” in Interspeech, Hyderabad, India . ISCA, 2018, pp. 3333–3337
2018
Cited alongside, same era.
S. Kim, S. Dalmia, and F. Metze, “Gated embeddings in end-to-end speech recognition for conversational-context fusion,” in ACL , Florence, Italy, Jul. 2019, pp. 1131–1141
2019
Later among the works it cites.
T. Mikolov and G. Zweig, “Context dependent recurrent neural network language model.” in SLT . IEEE, 2019, pp. 234–239
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
U. Khandelwal, H. He, P. Qi, and D. Jurafsky, “Sharp nearby, fuzzy far away: How neural language models use context,” in ACL , Jul. 2018, pp. 284–294
2018
Cited alongside, same era.
A. Sriram, H. Jun, S. Satheesh, and A. Coates, “Cold fusion: Training seq2seq models together with language models,” in Interspeech, Hyderabad, India . ISCA, 2018, pp. 387–391
2018
Cited alongside, same era.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Training language models for long-span cross-sentence evaluation,” in ASRU Singapore . IEEE, 2019, pp. 419–426
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT 2019 , Jun. 2019, pp. 4171–4186
2019
Cited alongside, same era.
S. Sukhbaatar, E. Grave, P. Bojanowski, and A. Joulin, “Adaptive attention span in transformers,” in ACL . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 331–335
2019
Cited alongside, same era.
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” in ACL , Florence, Italy, Jul. 2019, pp. 2978–2988
2019
Cited alongside, same era.
2019
Later among the works it cites.
D. Peskov, N. Clarke, J. Krone, B. Fodor, Y. Zhang, A. Youssef, and M. Diab, “Multi-domain goal-oriented dialogues (MultiDoGO): Strategies toward curating and annotating large scale dialogue data,” in Proc EMNLP-IJCNLP , 2019, pp. 4526–4536
2019
Later among the works it cites.
E. Hosseini-Asl, B. McCann, C.-S. Wu, S. Yavuz, and R. Socher, “A simple language model for task-oriented dialogue,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 20 179–20 191
2020
Later among the works it cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2020
2020
Later among the works it cites.
M. Kalimuthu, A. Mogadala, M. Mosbach, and D. Klakow, “Fusion models for improved image captioning,” in Pattern Recognition. ICPR International Workshops and Challenges , ser. Lecture Notes in Computer Science, vol. 12666, 2020, pp. 381–395
2020
Later among the works it cites.
M. Sunkara, S. Ronanki, D. Bekal, S. Bodapati, and K. Kirchhoff, “Multimodal semi-supervised learning framework for punctuation prediction in conversational speech,” in Interspeech 2020, Shanghai, China, 25-29 October 2020 . ISCA, 2020, pp. 4911–4915
2020
Later among the works it cites.
G. Sun, C. Zhang, and P. C. Woodland, “Transformer language models with lstm-based cross-utterance information representation,” 2021
2021
Closest in time.