Fetching the paper…
Reading the bibliography…
Autoregressive neural network models have been used successfully for sequence generation, feature extraction, and hypothesis scoring.
J. J. Rissanen, “Generalized Kraft inequality and arithmetic coding,”
1976
Earlier work this paper cites.
G. A. Frantz and R. H. Wiggins, “Design case history: Speak & Spell learns to talk,”
1982
Earlier work this paper cites.
I. H. Witten, R. M. Neal, and J. G. Cleary, “Arithmetic coding for data compression,”
1987
Earlier work this paper cites.
D. O’Shaughnessy, “Linear predictive coding,”
1988
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,”
1990
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,”
1991
Earlier work this paper cites.
J. Schmidhuber, “Neural sequence chunkers,” 1991
1991
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
K. E. Stanovich and R. F. West, “Individual differences in reasoning: Implications for the rationality debate?”
2000
Earlier work this paper cites.
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural probabilistic language model,”
2003
Earlier work this paper cites.
O. Chapelle, B. Schölkopf, and A. Zien, Eds.,
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
R. Levy, “Expectation-based syntactic comprehension,”
2008
Earlier work this paper cites.
D. Kahneman,
2011
Earlier work this paper cites.
——, “Integrating surprisal and uncertain-input models in online sentence comprehension: formal techniques and empirical results,” in
2011
Earlier work this paper cites.
Y. Huang and R. P. Rao, “Predictive coding,”
2011
Earlier work this paper cites.
H. Larochelle and I. Murray, “The neural autoregressive distribution estimator,” in
2011
Earlier work this paper cites.
I. F. Monsalve, S. L. Frank, and G. Vigliocco, “Lexical surprisal as a general predictor of reading time,” in
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Clark, “Whatever next? Predictive brains, situated agents, and the future of cognitive science,”
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: a literature survey,”
2014
Earlier work this paper cites.
D. Eigen, M. Ranzato, and I. Sutskever, “Learning factored representations in a deep mixture of experts,”
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,”
2014
Earlier work this paper cites.
W. Zaremba, I. Sutskever, and O. Vinyals, “Recurrent neural network regularization,”
2014
Earlier work this paper cites.
E. Park, D. Kim, S. Kim, Y.-D. Kim, G. Kim, S. Yoon, and S. Yoo, “Big/little deep neural network for ultra low power inference,” in
2015
Earlier work this paper cites.
S. Venkataramani, A. Raghunathan, J. Liu, and M. Shoaib, “Scalable-effort classifiers for energy-efficient machine learning,” in
2015
Earlier work this paper cites.
N. Léonard, “Distributed conditional computation,” 2015
2015
Cited alongside, same era.
A. M. Dai and Q. V. Le, “Semi-supervised sequence learning,” in
2015
Cited alongside, same era.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using LSTMs,” in
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and accurate recurrent neural network acoustic models for speech recognition,”
2015
Cited alongside, same era.
S. Teerapittayanon, B. McDanel, and H.-T. Kung, “Branchynet: Fast inference via early exiting from deep neural networks,” in
L. Liu and J. Deng, “Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution,” in
2018
Later among the works it cites.
C. Rosenbaum, T. Klinger, and M. Riemer, “Routing networks: Adaptive selection of non-linear functions for multi-task learning,”
2018
Later among the works it cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,”
2018
Later among the works it cites.
D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,”
2018
Later among the works it cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
S. Tan and K. C. Sim, “Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognition,” in
2016
Cited alongside, same era.
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup, “Conditional computation in neural networks for faster models,”
2016
Cited alongside, same era.
A. Graves, “Adaptive computation time for recurrent neural networks,”
2016
Cited alongside, same era.
K. M. Rocki, “Surprisal-driven feedback in recurrent networks,”
2016
Cited alongside, same era.
K. Rocki, T. Kornuta, and T. Maharaj, “Surprisal-driven zoneout,”
2016
Cited alongside, same era.
A. Gruenstein, R. Alvarez, C. Thornton, and M. Ghodrat, “A cascade architecture for keyword spotting on mobile devices,”
2017
Cited alongside, same era.
T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, “Neural ordinary differential equations,” in
2018
Later among the works it cites.
F. Lieder, A. Shenhav, S. Musslick, and T. L. Griffiths, “Rational metareasoning and the plasticity of cognitive control,”
2018
Later among the works it cites.
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,”
2019
Later among the works it cites.
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in NLP,”
2019
Later among the works it cites.
R. Tanno, K. Arulkumaran, D. C. Alexander, A. Criminisi, and A. Nori, “Adaptive neural trees,”
2019
Later among the works it cites.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser, “Universal transformers,”
2019
Later among the works it cites.
J. Ren, P. J. Liu, E. Fertig, J. Snoek, R. Poplin, M. Depristo, J. Dillon, and B. Lakshminarayanan, “Likelihood ratios for out-of-distribution detection,” in
2019
Later among the works it cites.
S. J. Mielke, R. Cotterell, K. Gorman, B. Roark, and J. Eisner, “What kind of language is hard to language-model?”
2019
Later among the works it cites.
H. He, N. Peng, and P. Liang, “Pun generation with surprise,”
2019
Later among the works it cites.
T. Alpay, F. Abawi, and S. Wermter, “Preserving activations in recurrent neural networks based on surprisal,”
2019
Later among the works it cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,”
2019
Later among the works it cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,”
2019
Later among the works it cites.
F. S. Fard and T. Trappenberg, “A novel model for arbitration between planning and habitual control systems,”
2019
Later among the works it cites.
M. Peters, S. Ruder, and N. A. Smith, “To tune or not to tune? Adapting pretrained representations to diverse tasks,”
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga
2019
Later among the works it cites.
O. Sharir, B. Peleg, and Y. Shoham, “The Cost of Training NLP Models: A Concise Overview,”
2020
Closest in time.
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference,”
2020
Closest in time.
S. Scardapane, M. Scarpiniti, E. Baccarelli, and A. Uncini, “Why should we add early exits to neural networks?” 2020
2020
Closest in time.
J. Dean, “The deep learning revolution and its implications for computer architecture and chip design,” in
2020
Closest in time.
2020
Closest in time.
A. Goyal, Y. Bengio, and M. B. S. Levine, “The variational bandwidth bottleneck: Stochastic evaluation on an information budget,”
2020
Closest in time.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2020
2020
Closest in time.