Fetching the paper…
Reading the bibliography…
In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2006
Earlier work this paper cites.
Cluster-former: Clustering-based sparse transformer for long-range dependency encoding
Wang, S., Zhou, L., Gan, Z., Chen, Y.-C., Fang, Y., Sun, S., Cheng, Y., and Liu, J · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Parikh, A., Täckström, O., Das, D., and Uszkoreit, J · 2016
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Yu, F. and Koltun, V · 2016
Earlier work this paper cites.
Reading wikipedia to answer open-domain questions
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Simple and effective multi-paragraph reading comprehension
Clark, C. and Gardner, M · 2018
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N · 2018
Earlier work this paper cites.
Document-level neural machine translation with hierarchical attention networks
Miculicich, L., Ram, D., Pappas, N., and Henderson, J · 2018
Earlier work this paper cites.
A bert baseline for the natural questions
Alberti, C., Lee, K., and Collins, M · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Star-transformer
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Cited alongside, same era.
Hierarchical transformers for multi-document summarization
Liu, Y. and Lapata, M · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Deberta: Decoding-enhanced bert with disentangled attention, 2020
He, P., Liu, X., Gao, J., and Chen, W · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Later among the works it cites.
Rikinet: Reading wikipedia pages for natural question answering
Liu, D., Gong, Y., Fu, J., Yan, Y., Chen, J., Jiang, D., Lv, J., and Duan, N · 2020
Later among the works it cites.
On extractive and abstractive neural document summarization with transformer language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, F., Fan, A., Baevski, A., Dauphin, Y. N., and Auli, M · 2019
Cited alongside, same era.
Bp-transformer: Modelling long-range context via binary partitioning
Ye, Z., Guo, Q., Gan, Q., Qiu, X., and Zhang, Z · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
Improving multilingual models with language-clustered vocabularies
Chung, H. W., Garrette, D., Tan, K. C., and Riesa, J · 2020
Cited alongside, same era.
Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J · 2020
Cited alongside, same era.
Pilault, J., Li, R., Subramanian, S., and Pal, C · 2020
Later among the works it cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Qi, W., Yan, Y., Gong, Y., Liu, D., Duan, N., Chen, J., Zhang, R., and Zhou, M · 2020
Later among the works it cites.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, W.-t., Wang, S., and Tang, J · 2020
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Later among the works it cites.
Synthesizer: Rethinking self-attention in transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontañón, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Zhang, J., Zhao, Y., Saleh, M., and Liu, P · 2020
Later among the works it cites.