Fetching the paper…
Reading the bibliography…
State-of-the-art neural language models (LMs) represented by Transformers are highly complex.
“Improved backing-off for m-gram language modeling,”
Reinhard Kneser and Hermann Ney, · 1995
Earlier work this paper cites.
“Ensemble learning in bayesian neural networks,”
David Barber and Christopher M Bishop, · 1998
Earlier work this paper cites.
“Extensions of recurrent neural network language model,”
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur, · 2011
Earlier work this paper cites.
“Practical variational inference for neural networks,”
Alex Graves, · 2011
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Auto-encoding variational bayes,”
Diederik P Kingma and Max Welling, · 2013
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,”
Pawel Swietojanski and Steve Renals, · 2014
Earlier work this paper cites.
“Recurrent neural network based language model,”
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernock, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Dropout as a bayesian approximation: Representing model uncertainty in deep learning,”
Yarin Gal and Zoubin Ghahramani, · 2015
Earlier work this paper cites.
“Variational dropout and the local reparameterization trick,”
Diederik P Kingma, Tim Salimans, and Max Welling, · 2015
Cited alongside, same era.
“Long short-term memory-networks for machine reading,”
Jianpeng Cheng, Li Dong, and Mirella Lapata, · 2016
Cited alongside, same era.
“A decomposable attention model for natural language inference,”
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit, · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, X. Zhang, Shaoqing Ren, and J. Sun, · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Cited alongside, same era.
“Bayesian recurrent neural network for language modeling,”
Jen-Tzung Chien and Yuan Chu Ku, · 2016
“Convolutional sequence to sequence learning,”
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin, · 2017
Later among the works it cites.
“Improving language understanding by generative pre-training,” 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Later among the works it cites.
“Semi-orthogonal low-rank matrix factorization for deep neural networks.,”
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohammadi, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Language Modeling with Deep Transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Gaussian process lstm recurrent neural network language models for speech recognition,”
Max W. Y. Lam, Xie Chen, Shoukang Hu, Jianwei Yu, and Helen M. Meng, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Bridging nonlinearities and stochastic regularizers with gaussian error linear units,”
Dan Hendrycks and Kevin Gimpel, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for asr based on lattice-free mmi,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“A structured self-attentive sentence embedding,”
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio, · 2017
Cited alongside, same era.
“Blt: Exact bayesian inference with distribution transformers,”
Charles Yuan and Jan Hoffmann,
Cited in the paper.
Jianwei Yu, Max W. Y. Lam, Shoukang Hu, Xixin Wu, and Helen M. Meng, · 2019
Later among the works it cites.
“Bayesian layers: A module for neural network uncertainty,”
Dustin Tran, Mike Dusenberry, Mark van der Wilk, and Danijar Hafner, · 2019
Later among the works it cites.
“An empirical study of transformer-based neural language model adaptation,”
Ke Li, Zhe Liu, Tianxing He, Hongzhao Huang, and Sanjeev Khudanpur, · 2020
Later among the works it cites.
“Development of the cuhk elderly speech recognition system for neurocognitive disorder detection using the dementiabank corpus,”
Zi YE, Shoukang Hu, Jinchao Li, Xurong Xie, Mengzhe Geng, Jianwei Yu, Junhao Xu, Boyang Xue, Shansong Liu, Xunying Liu, and Helen Meng, · 2021
Closest in time.