Fetching the paper…
Reading the bibliography…
Modeling long texts has been an essential technique in the field of natural language processing (NLP).
Signature verification using a” siamese” time delay neural network
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger, and Roopak Shah · 1993
Earlier work this paper cites.
Textrank: Bringing order into text
Rada Mihalcea and Paul Tarau · 2004
Earlier work this paper cites.
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Summary level training of sentence rewriting for abstractive summarization
Sanghwan Bae, Taeuk Kim, Jihoon Kim, and Sang-goo Lee · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Semantic text matching for long-form documents
Jyun-Yu Jiang, Mingyang Zhang, Cheng Li, Michael Bendersky, Nadav Golbandi, and Marc Najork · 2019
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Hierarchical transformers for long document classification
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba, Yishay Carmiel, and Najim Dehak · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Long-length legal document classification
Lulu Wan, George Papageorgiou, Michael Seddon, and Mirko Bernardoni · 2019
Earlier work this paper cites.
Multi-passage bert: A globally normalized bert model for open-domain question answering
Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, and Bing Xiang · 2019
Earlier work this paper cites.
HIBERT: Document level pre-training of hierarchical bidirectional transformers for document summarization
Xingxing Zhang, Furu Wei, and Ming Zhou · 2019
Earlier work this paper cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy J. Colwell, and Adrian Weller · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Enhancing extractive text summarization with topic-aware graph neural networks
Peng Cui, Le Hu, and Yuanchao Liu · 2020
Earlier work this paper cites.
Cogltx: Applying bert to long texts
Ming Ding, Chang Zhou, Hongxia Yang, and Jie Tang · 2020
Earlier work this paper cites.
Recurrent chunking mechanisms for long-text machine reading comprehension
Hongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen, and Dong Yu · 2020
Cited alongside, same era.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Edouard Grave · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Parade: Passage representation aggregation for document reranking
Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun · 2020
Cited alongside, same era.
Hierarchical learning for generation with long source sequences
Tobias Rohde, Xiaoxia Wu, and Yinhan Liu · 2021
Later among the works it cites.
RoR: Read-over-read for long document machine reading comprehension
Jing Zhao, Junwei Bao, Yifan Wang, Yongwei Zhou, Youzheng Wu, Xiaodong He, and Bowen Zhou · 2021
Later among the works it cites.
HIBRIDS: Attention with hierarchical biases for structure-aware long document summarization
Shuyang Cao and Lu Wang · 2022
Later among the works it cites.
An exploration of hierarchical attention transformers for efficient long document classification
Ilias Chalkidis, Xiang Dai, Manos Fergadiotis, Prodromos Malakasiotis, and Desmond Elliott · 2022
Later among the works it cites.
Toward unifying text segmentation and long document summarization
Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu, and Dong Yu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On extractive and abstractive neural document summarization with transformer language models
Jonathan Pilault, Raymond Li, Sandeep Subramanian, and Chris Pal · 2020
Cited alongside, same era.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang · 2020
Cited alongside, same era.
Sparse sinkhorn attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Memformer: The memory-augmented transformer
Qingyang Wu, Zhenzhong Lan, Jing Gu, and Zhou Yu · 2020
Cited alongside, same era.
Systematically exploring redundancy reduction in summarizing long documents
Wen Xiao and Giuseppe Carenini · 2020
Cited alongside, same era.
Beyond 512 tokens: Siamese multi-depth transformer-based hierarchical encoder for long-form document matching
Liu Yang, Mingyang Zhang, Cheng Li, Michael Bendersky, and Marc Najork · 2020
Cited alongside, same era.
Later among the works it cites.
Multi graph neural network for extractive long document summarization
Xuan-Dung Doan, Le-Minh Nguyen, and Khac-Hoai Nam Bui · 2022
Later among the works it cites.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang · 2022
Later among the works it cites.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant · 2022
Later among the works it cites.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan · 2022
Later among the works it cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2022
Later among the works it cites.
Leveraging locality in abstractive text summarization
Yixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb, Chenguang Zhu, Ahmed H Awadallah, and Dragomir Radev · 2022
Later among the works it cites.
DYLE: Dynamic latent extraction for abstractive long-input summarization
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Awadallah, and Dragomir Radev · 2022
Later among the works it cites.
Semantic self-segmentation for abstractive summarization of long documents in low-resource regimes
Gianluca Moro and Luca Ragazzi · 2022
Later among the works it cites.
Yuxiang Nie, Heyan Huang, Wei Wei, and Xian-Ling Mao · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Long document summarization with top-down and bottom-up inference
Bo Pang, Erik Nijkamp, Wojciech Kryściński, Silvio Savarese, Yingbo Zhou, and Caiming Xiong · 2022
Later among the works it cites.
Efficient classification of long documents using transformers
Hyunji Park, Yogarshi Vyas, and Kashif Shah · 2022
Later among the works it cites.
HeterGraphLongSum: Heterogeneous graph neural network with passage aggregation for extractive long document summarization
Tuan-Anh Phan, Ngoc-Dung Ngoc Nguyen, and Khac-Hoai Nam Bui · 2022
Later among the works it cites.
Investigating efficiently extending transformers for long input summarization
Jason Phang, Yao Zhao, and Peter J Liu · 2022
Later among the works it cites.
Two-stage movie script summarization: An efficient method for low-resource long document summarization
Dongqi Pu, Xudong Hong, Pin-Jie Lin, Ernie Chang, and Vera Demberg · 2022
Later among the works it cites.
Parallel context windows improve in-context learning of large language models
Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2022
Later among the works it cites.
HiStruct+: Improving extractive text summarization with hierarchical structure information
Qian Ruan, Malte Ostendorff, and Georg Rehm · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Later among the works it cites.
Hegel: Hypergraph transformer for long document summarization
Haopeng Zhang, Xiao Liu, and Jiawei Zhang · 2022
Later among the works it cites.
Efficient long sequence modeling via state space augmented transformer
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao · 2022
Later among the works it cites.