Fetching the paper…
Reading the bibliography…
We investigate multi-scale transformer language models that learn representations of text at multiple scales, and present three different architectures that have an inductive bias to handle the hierarchical nature of language.
The laplacian pyramid as a compact image code
Burt, P. and Adelson, E · 1983
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J · 1992
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
El Hihi, S. and Bengio, Y · 1996
Earlier work this paper cites.
The uncrowded window of object recognition
Pelli, D. G. and Tillman, K. A · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Metamers of the ventral stream
Freeman, J. and Simoncelli, E. P · 2011
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Heafield, K · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J · 2014
Earlier work this paper cites.
Deep generative image models using a laplacian pyramid of adversarial networks
Denton, E. L., Chintala, S., Fergus, R., et al · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
A hierarchical recurrent encoder-decoder for generative context-aware query suggestion
Sordoni, A., Bengio, Y., Vahabi, H., Lioma, C., Grue Simonsen, J., and Nie, J.-Y · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Later among the works it cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Garg, V. K., Dhillon, I. S., and Yu, H.-F · 2019
Later among the works it cites.
Sample efficient text summarization using a single pre-trained transformer
Khandelwal, U., Clark, K., Jurafsky, D., and Kaiser, L · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2018
Cited alongside, same era.
Eval all, trust a few, do wrong to none: Comparing sentence generation models
Cífka, O., Severyn, A., Alfonseca, E., and Filippova, K · 2018
Cited alongside, same era.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Cited alongside, same era.
Sharp nearby, fuzzy far away: How neural language models use context
Khandelwal, U., He, H., Qi, P., and Jurafsky, D · 2018
Cited alongside, same era.
Generating wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N · 2018
Cited alongside, same era.
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2019
Later among the works it cites.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
Later among the works it cites.
Hierarchical transformers for multi-document summarization
Liu, Y. and Lapata, M · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 2019
Later among the works it cites.
Do neural dialog systems use the conversation history effectively? an empirical study
Sankar, C., Subramanian, S., Pal, C., Chandar, S., and Bengio, Y · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
Image content is more important than bouma’s law for scene metamers
Wallis, T. S., Funke, C. M., Ecker, A. S., Gatys, L. A., Wichmann, F. A., and Bethge, M · 2019
Later among the works it cites.
Neural text generation with unlikelihood training
Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J · 2019
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Closest in time.