Fetching the paper…
Reading the bibliography…
The multi-head self-attention mechanism of the transformer model has been thoroughly investigated recently.
Docbert: Bert for document classification
Ashutosh Adhikari, Achyudh Ram, Raphael Tang, and Jimmy Lin. 2019 · 1904
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Sparck Jones. 1972 · 1972
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
The new york times annotated corpus
Evan Sandhaus. 2008 · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Summarunner: A recurrent neural network based sequence model for extractive summarization of documents
Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017 · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Earlier work this paper cites.
Text segmentation as a supervised learning task
Omri Koshorek, Adir Cohen, Noam Mor, Michael Rotman, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2018 · 2018
Earlier work this paper cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Earlier work this paper cites.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Earlier work this paper cites.
SECTOR: A neural model for coherent topic segmentation and classification
Sebastian Arnold, Rudolf Schneider, Philippe Cudré-Mauroux, Felix A. Gers, and Alexander Löser. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019 · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. 2019 · 2019
Earlier work this paper cites.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
Discourse-aware semantic self-attention for narrative reading comprehension
Todor Mihaylov and Anette Frank. 2019 · 2019
Cited alongside, same era.
Importance estimation for neural network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019 · 2019
Cited alongside, same era.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Do we really need that many parameters in transformer for extractive summarization? discourse can help !
Wen Xiao, Patrick Huber, and Giuseppe Carenini. 2020 · 2020
Later among the works it cites.
Improving context modeling in neural topic segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Definitions, methods, and applications in interpretable machine learning
W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
How well do NLI models capture verb veridicality?
Alexis Ross and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Cited alongside, same era.
Linzi Xing, Brad Hackinen, Giuseppe Carenini, and Francesco Trebbi. 2020 · 2020
Later among the works it cites.
Discourse-aware neural extractive text summarization
Jiacheng Xu, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
How does BERT’s attention change when you fine-tune? an analysis methodology and a case study in negation scope
Yiyun Zhao and Steven Bethard. 2020 · 2020
Later among the works it cites.
Syntax-BERT: Improving pre-trained transformers with syntax trees
Jiangang Bai, Yujing Wang, Yiren Chen, Yaming Yang, Jing Bai, Jing Yu, and Yunhai Tong. 2021 · 2021
Closest in time.
On attention redundancy: A comprehensive study
Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, and Kenneth Church. 2021 · 2021
Closest in time.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021 · 2021
Closest in time.
Bird’s eye: Probing for linguistic graph structures with a simple information-theoretic approach
Yifan Hou and Mrinmaya Sachan. 2021 · 2021
Closest in time.
R2D2: Recursive transformer based on differentiable tree for interpretable hierarchical language modeling
Xiang Hu, Haitao Mi, Zujie Wen, Yafang Wang, Yi Su, Jing Zheng, and Gerard de Melo. 2021 · 2021
Closest in time.
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. 2021 · 2021
Closest in time.
T3-vis: visual analytic for training and fine-tuning transformers in NLP
Raymond Li, Wen Xiao, Lanjun Wang, Hyeju Jang, and Giuseppe Carenini. 2021 · 2021
Closest in time.
Transformer over pre-trained transformer for neural text segmentation with enhanced topic coherence
Kelvin Lo, Yuan Jin, Weicong Tan, Ming Liu, Lan Du, and Wray Buntine. 2021 · 2021
Closest in time.
Sparsebert: Rethinking the importance analysis in self-attention
Han Shi, Jiahui Gao, Xiaozhe Ren, Hang Xu, Xiaodan Liang, Zhenguo Li, and James Tin-Yau Kwok. 2021 · 2021
Closest in time.
Synthesizer: Rethinking self-attention for transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2021 · 2021
Closest in time.
Predicting discourse trees from transformer-based neural summarizers
Wen Xiao, Patrick Huber, and Giuseppe Carenini. 2021 · 2021
Closest in time.
Towards understanding large-scale discourse structures in pre-trained and fine-tuned language models
Patrick Huber and Giuseppe Carenini. 2022 · 2022
Closest in time.
HiStruct+: Improving extractive text summarization with hierarchical structure information
Qian Ruan, Malte Ostendorff, and Georg Rehm. 2022 · 2022
Closest in time.
Learning adaptive axis attentions in fine-tuning: Beyond fixed sparse attention patterns
Zihan Wang, Jiuxiang Gu, Jason Kuen, Handong Zhao, Vlad Morariu, Ruiyi Zhang, Ani Nenkova, Tong Sun, and Jingbo Shang. 2022 · 2022
Closest in time.
Metaphors in pre-trained language models: Probing and generalization across datasets and languages
Ehsan Aghazadeh, Mohsen Fayyaz, and Yadollah Yaghoobzadeh. 2022 · 2050
Closest in time.