Fetching the paper…
Reading the bibliography…
Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions.
Scripts, plans, goals, and understanding: an inquiry into human knowledge structures
Roger C Schank and Robert P Abelson. 1977 · 1977
Earlier work this paper cites.
Coherence and coreference
Jerry R Hobbs. 1979 · 1979
Earlier work this paper cites.
Learning schemata for natural language processing
Raymond J Mooney and Gerald DeJong. 1985 · 1985
Earlier work this paper cites.
Argument structure
Jane Grimshaw. 1990 · 1990
Earlier work this paper cites.
Centering: A framework for modeling the local coherence of discourse
Barbara J. Grosz, Aravind K. Joshi, and Scott Weinstein. 1995 · 1995
Earlier work this paper cites.
A maximum entropy approach to adaptive statistical language modeling
Roni Rosenfeld. 1996 · 1996
Earlier work this paper cites.
Narrative analysis: Oral versions of personal experience
William Labov and Joshua Waletzky. 1997 · 1997
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Document context language models
Yangfeng Ji, Trevor Cohn, Lingpeng Kong, Chris Dyer, and Jacob Eisenstein. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2015 · 2015
Earlier work this paper cites.
Larger-context language modelling with recurrent neural network
Tian Wang and Kyunghyun Cho. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
Prediction with a short memory
Vatsal Sharan, Sham Kakade, Percy Liang, and Gregory Valiant. 2018 · 2018
Cited alongside, same era.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Improving the transformer translation model with document-level context
Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, Min Zhang, and Yang Liu. 2018 · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Context analysis for pre-trained masked language models
Yi-An Lai, Garima Lalwani, and Yi Zhang. 2020 · 2020
Later among the works it cites.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A. Smith, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi. 2020 · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Later among the works it cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2020
Cited alongside, same era.
Local self-attention over long text for efficient document retrieval
Sebastian Hofstätter, Hamed Zamani, Bhaskar Mitra, Nick Craswell, and Allan Hanbury. 2020 · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2020
Later among the works it cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Later among the works it cites.
Lite transformer with long-short range attention
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
Long-short term masking transformer: A simple but effective baseline for document-level neural machine translation
Pei Zhang, Boxing Chen, Niyu Ge, and Kai Fan. 2020 · 2020
Later among the works it cites.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. 2021 · 2021
Closest in time.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021 · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Closest in time.
Synthesizer: Rethinking self-attention for transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2021 · 2021
Closest in time.