Fetching the paper…
Reading the bibliography…
Transformers-based models, such as BERT, have been one of the most successful deep learning models for NLP.
Multi-hop question answering via reasoning chains
J. Chen, S.-t. Lin, and G. Durrett · 1910
Earlier work this paper cites.
Distilling the knowledge of bert for text generation
Y.-C. Chen, Z. Gan, Y. Cheng, J. Liu, and J. Liu · 1911
Earlier work this paper cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu · 1912
Earlier work this paper cites.
Long-range correlation properties of coding and noncoding dna sequences: Genbank analysis
S. Buldyrev, A. Goldberger, S. Havlin, R. Mantegna, M. Matsa, C.-K. Peng, M. Simons, and H. Stanley · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Collective dynamics of ‘small-world’networks
D. J. Watts and S. H. Strogatz · 1998
Earlier work this paper cites.
The average distances in random graphs with given expected degrees
F. Chung and L. Lu · 2002
Earlier work this paper cites.
The human genome browser at ucsc
W. J. Kent, C. W. Sugnet, T. S. Furey, K. M. Roskin, T. H. Pringle, A. M. Zahler, and D. Haussler · 2002
Earlier work this paper cites.
What makes a good answer? the role of context in question answering
J. Lin, D. Quan, V. Sinha, K. Bakshi, D. Huynh, B. Katz, and D. R. Karger · 2003
Earlier work this paper cites.
Lexrank: Graph-based lexical centrality as salience in text summarization
G. Erkan and D. R. Radev · 2004
Earlier work this paper cites.
The impact of frequency on summarization
A. Nenkova and L. Vanderwende · 2005
Earlier work this paper cites.
A new algorithm for optimal 2-constraint satisfaction and its implications
R. Williams · 2005
Earlier work this paper cites.
Expander graphs and their applications
S. Hoory, N. Linial, and A. Wigderson · 2006
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Spectral sparsification of graphs
D. A. Spielman and S.-H. Teng · 2011
Earlier work this paper cites.
Segmenting dna sequence into words based on statistical language model
W. Liang · 2012
Earlier work this paper cites.
Epd and epdnew, high-quality promoter resources in the next-generation sequencing era
R. Dreos, G. Ambrosini, R. Cavin Périer, and P. Bucher · 2013
Earlier work this paper cites.
Consequences of faster alignment of sequences
A. Abboud, V. V. Williams, and O. Weimann · 2014
Earlier work this paper cites.
Enhanced regulatory sequence prediction using gapped k-mer features
M. Ghandi, D. Lee, M. Mohammad-Noori, and M. A. Beer · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Tight hardness results for lcs and other sequence similarity measures
A. Abboud, A. Backurs, and V. V. Williams · 2015
Earlier work this paper cites.
Edit distance cannot be computed in strongly subquadratic time (unless seth is false)
A. Backurs and P. Indyk · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
J. Zhou and O. G. Troyanskaya · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Role of non-coding sequence variants in cancer
E. Khurana, Y. Fu, D. Chakravarty, F. Demichelis, M. A. Rubin, and M. Gerstein · 2016
Earlier work this paper cites.
Simple and effective multi-paragraph reading comprehension
C. Clark and M. Gardner · 2017
Earlier work this paper cites.
Histone marks in the ‘driver’s seat’: functional roles in steering the transcription cycle
L. A. Gates, C. E. Foulds, and B. W. O’Malley · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Earlier work this paper cites.
Gpu kernels for block-sparse weights
S. Gray, A. Radford, and D. P. Kingma · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Identifying sigma70 promoters with novel pseudo nucleotide composition
H. Lin, Z.-Y. Liang, H. Tang, and W. Chen · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
A. See, P. J. Liu, and C. D. Manning · 2017
Earlier work this paper cites.
Lecture Notes for Boston University MA 882 Spring 2017 , 2017 (accessed June 3, 2020)
D. Sussman · 2017
Earlier work this paper cites.
Recognition of prokaryotic and eukaryotic promoters using convolutional deep learning neural networks
R. K. Umarov and V. V. Solovyev · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
S. Wiseman, S. M. Shieber, and A. M. Rush · 2017
Cited alongside, same era.
Exploiting sequence-based features for predicting enhancer–promoter interactions
Y. Yang, R. Zhang, S. Singh, and J. Ma · 2017
Cited alongside, same era.
Promoterpredict: sequence-based modelling of escherichia coli σ \sigma 70 promoter strength yields logarithmic dependence between promoter strength and sequence
R. Bharanikumar, K. A. R. Premkumar, and A. Palaniappan · 2018
Cited alongside, same era.
A discourse-aware attention model for abstractive summarization of long documents
A. Cohan, F. Dernoncourt, D. S. Kim, T. Bui, S. Kim, W. Chang, and N. Goharian · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Camembert: a tasty french language model
L. Martin, B. Muller, P. J. O. Suárez, Y. Dupont, L. Romary, É. V. de la Clergerie, D. Seddah, and B. Sagot · 2019
Later among the works it cites.
Leveraging bert for extractive text summarization on lectures
D. Miller · 2019
Later among the works it cites.
Adapting pretrained language models for long document classification
M. L. Olson, L. Zhang, and C.-N. Yu · 2019
Later among the works it cites.
Deepromoter: Robust promoter predictor using deep learning
M. Oubounyt, Z. Louadi, H. Tayara, and K. T. Chong · 2019
Later among the works it cites.
On the turing completeness of modern neural network architectures
J. Pérez, J. Marinković, and P. Barceló · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Bottom-up abstractive summarization
S. Gehrmann, Y. Deng, and A. M. Rush · 2018
Cited alongside, same era.
Distribution of shortest path lengths in subcritical erdős-rényi networks
E. Katzav, O. Biham, and A. K. Hartmann · 2018
Cited alongside, same era.
Abstractive summarization of reddit posts with multi-level memory networks
B. Kim, H. Kim, and G. Kim · 2018
Cited alongside, same era.
T. Kudo and J. Richardson · 2018
Cited alongside, same era.
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Cited alongside, same era.
Blockwise self-attention for long document understanding
J. Qiu, H. Ma, O. Levy, S. W.-t. Yih, S. Wang, and J. Tang · 2019
Later among the works it cites.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Later among the works it cites.
Leveraging pre-trained checkpoints for sequence generation tasks
S. Rothe, S. Narayan, and A. Severyn · 2019
Later among the works it cites.
Bigpatent: A large-scale dataset for abstractive and coherent summarization
E. Sharma, C. Li, and L. Wang · 2019
Later among the works it cites.
On extractive and abstractive neural document summarization with transformer language models
S. Subramanian, R. Li, J. Pilault, and C. Pal · 2019
Later among the works it cites.
Adaptive attention span in transformers
S. Sukhbaatar, E. Grave, P. Bojanowski, and A. Joulin · 2019
Later among the works it cites.
Utilizing bert for aspect-based sentiment analysis via constructing auxiliary sentence
C. Sun, L. Huang, and X. Qiu · 2019
Later among the works it cites.
Viraminer: Deep learning on raw dna sequences for identifying viral genomes in human samples
A. Tampuu, Z. Bzhalava, J. Dillner, and R. Vicente · 2019
Later among the works it cites.
Sentiment classification using document embeddings trained with cosine similarity
T. Thongtan and T. Phienthrakul · 2019
Later among the works it cites.
Multi-passage bert: A globally normalized bert model for open-domain question answering
Z. Wang, P. Ng, X. Ma, R. Nallapati, and B. Xiang · 2019
Later among the works it cites.
ipsw (2l)-pseknc: A two-layer predictor for identifying promoters and their strength by hybrid features via pseudo k-tuple nucleotide composition
X. Xiao, Z.-C. Xu, W.-R. Qiu, P. Wang, H.-T. Ge, and K.-C. Chou · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Later among the works it cites.
Balanced sparsity for efficient dnn inference on gpu
Z. Yao, S. Cao, W. Xiao, C. Zhang, and L. Nie · 2019
Later among the works it cites.
Bp-transformer: Modelling long-range context via binary partitioning
Z. Ye, Q. Guo, Q. Gan, X. Qiu, and Z. Zhang · 2019
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
C. Yun, S. Bhojanapalli, A. S. Rawat, S. J. Reddi, and S. Kumar · 2019
Later among the works it cites.
Etc: Encoding long and structured data in transformers
J. Ainslie, S. Ontanon, C. Alberti, P. Pham, A. Ravula, and S. Sanghai · 2020
Closest in time.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Closest in time.
Spectral radii of sparse random matrices
F. Benaych-Georges, C. Bordenave, A. Knowles, et al · 2020
Closest in time.
A divide-and-conquer approach to the summarization of academic articles
A. Gidiotis and G. Tsoumakas · 2020
Closest in time.
ReflectionNet , 2020 (accessed June 3, 2020)
M. Gong · 2020
Closest in time.
Realm: Retrieval-augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang · 2020
Closest in time.
Leveraging passage retrieval with generative models for open domain question answering
G. Izacard and E. Grave · 2020
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy · 2020
Closest in time.
Data augmentation using pre-trained transformer models
V. Kumar, A. Choudhary, and E. Cho · 2020
Closest in time.
Patent classification by fine-tuning bert language model
J.-S. Lee and J. Hsiang · 2020
Closest in time.
Methylnet: an automated and modular deep learning approach for dna methylation analysis
J. J. Levy, A. J. Titus, C. L. Petersen, Y. Chen, L. A. Salas, and B. C. Christensen · 2020
Closest in time.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Closest in time.
Rikinet: Reading wikipedia pages for natural question answering
D. Liu, Y. Gong, J. Fu, Y. Yan, J. Chen, D. Jiang, J. Lv, and N. Duan · 2020
Closest in time.
Multi-hop reading comprehension across documents with path-based graph convolutional network
Z. Tang, Y. Shen, X. Ma, W. Xu, J. Yu, and W. Lu · 2020
Closest in time.
o ( n ) o(n) connections are expressive enough: Universal approximability of sparse transformers
C. Yun, Y.-W. Chang, S. Bhojanapalli, A. S. Rawat, S. J. Reddi, and S. Kumar · 2020
Closest in time.