Fetching the paper…
Reading the bibliography…
Transformer-based language models benefit from conditioning on contexts of hundreds to thousands of previous tokens.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 1905
Earlier work this paper cites.
Do neural dialog systems use the conversation history effectively? an empirical study
Chinnadhurai Sankar, Sandeep Subramanian, Christopher Pal, Sarath Chandar, and Yoshua Bengio. 2019 · 1906
Earlier work this paper cites.
A mathematical theory of communication
Claude E Shannon. 1948 · 1948
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
An estimate of an upper bound for the entropy of english
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, Jennifer C Lai, and Robert L Mercer. 1992 · 1992
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney. 1995 · 1995
Earlier work this paper cites.
A bit of progress in language modeling
Joshua T Goodman. 2001 · 2001
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Modeling local coherence: An entity-based approach
Regina Barzilay and Mirella Lapata. 2008 · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
The CMU-EBMT machine translation system
Ralf D Brown. 2011 · 2011
Earlier work this paper cites.
Thang M Pham, Trung Bui, Long Mai, and Anh Nguyen. 2020 · 2012
Cited alongside, same era.
Tracking the world state with recurrent entity networks
Mikael Henaff, Jason Weston, Arthur Szlam, Antoine Bordes, and Yann LeCun. 2016 · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Language GANs falling short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin. 2020 · 2020
Later among the works it cites.
Curious case of language generation evaluation metrics: A cautionary tale
Ozan Caglayan, Pranava Madhyastha, and Lucia Specia. 2020 · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Unifying human and statistical evaluation for natural language generation
Tatsunori Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Composition is the core driver of the language-selective network
F. Mollica, Matthew Siegelman, Evgeniia Diachek, S. Piantadosi, Zachary Mineroff, Richard Futrell, Hope H. Kean, Peng Qian, and E. Fedorenko. 2020 · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020 · 2020
Later among the works it cites.
Rissanen data analysis: Examining dataset characteristics via description length
Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021 · 2021
Closest in time.