Fetching the paper…
Reading the bibliography…
This paper presents the Ensemble Nucleotide Byte-level Encoder-Decoder (ENBED) foundation model, analyzing DNA sequences at byte-level precision with an encoder-decoder Transformer architecture.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, July 2020
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 1910
Earlier work this paper cites.
On statistical modeling of sequencing noise in high depth data to assess tumor evolution
R. Rabadan, G. Bhanot, S. Marsilio, N. Chiorazzi, L. Pasqualucci, and H. Khiabanian · 1945
Earlier work this paper cites.
The influenza virus resource at the national center for biotechnology information
Y. Bao, P. Bolotov, D. Dernovoy, B. Kiryutin, L. Zaslavsky, T. Tatusova, J. Ostell, and D. Lipman · 2008
Earlier work this paper cites.
Gencode: The reference human genome annotation for the encode project
J. Harrow et. al · 2012
Earlier work this paper cites.
A global reference for human genetic variation
. G. P. Consortium · 2015
Earlier work this paper cites.
Reference sequence (refseq) database at ncbi: current status, taxonomic expansion, and functional annotation
N. A. O’Leary, M. W. Wright, J. R. Brister, et al · 2016
Earlier work this paper cites.
Recognition of prokaryotic and eukaryotic promoters using convolutional deep learning neural networks
R. K. Umarov and V. V. Solovyev · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, et al · 2018
Earlier work this paper cites.
Enhancer identification using transfer and adversarial deep learning of dna sequences
D. Cohn, O. Zuk, and T. Kaplan · 2018
Earlier work this paper cites.
Mutagan: A sequence-to-sequence gan framework to predict mutations of evolving protein populations
D. S. Berman, C. Howser, T. Mehoke, and J. D. Evans · 2020
Earlier work this paper cites.
Bertology meets biology: Interpreting attention in protein language models
J. Vig, A. Madani, L. R. Varshney, et al · 2020
Earlier work this paper cites.
Effective gene expression prediction from sequence by integrating long-range interactions
Ž. Avsec, V. Agarwal, D. Visentin, et al · 2021
Earlier work this paper cites.
ProtTrans: Toward understanding the language of life through self-supervised learning
A. Elnaggar, M. Heinzinger, C. Dallago, et al · 2021
Earlier work this paper cites.
Exploring the limits of out-of-distribution detection
S. Fort, J. J. Ren, and B. Lakshminarayanan · 2021
Earlier work this paper cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
A. Rives, J. Meier, T. Sercu, S. Goyal, et al · 2021
Cited alongside, same era.
Prediction of rna–protein interactions using a nucleotide language model
K. Yamada and M. Hamada · 2021
Cited alongside, same era.
A method for multiple-sequence-alignment-free protein structure prediction using a protein language model
X. Fang, F. Wang, L. Liu, J. He, D. Lin, Y. Xiang, X. Zhang, H.-H. Wu, H. Li, and L. Song · 2022
Cited alongside, same era.
A deep learning framework for enhancer prediction using word embedding and sequence generation
Q. Geng, R. Yang, and L. Zhang · 2022
Cited alongside, same era.
Genomic benchmarks: a collection of datasets for genomic sequence classification
K. Grevsova, V. Martinek, D. Cechak, P. Simecek, and P. Alexiou · 2022
Cited alongside, same era.
The nucleotide transformer: Building and evaluating robust foundation models for human genomics
H. Dalla-Torre, L. Gonzalez, J. M. Revilla, N. L. Carranza, A. H. Grzywaczewski, F. Oteri, C. Dallago, E. Trop, H. Sirelkhatim, G. Richard, M. J. Skwark, K. Beguir, M. Lopez, and T. Pierrot · 2023
Closest in time.
Effect of tokenization on transformers for biological sequences
E. Dotan, G. Jaschek, T. Pupko, and Y. Belinkov · 2023
Closest in time.
GENA-LM: A Family of Open-Source Foundational Models for Long DNA Sequences
V. Fishman, Y. Kuratov, M. Petrov, A. Shmelev, D. Shepelin, N. Chekanov, O. Kardymon, and M. Burtsev · 2023
Closest in time.
Decoder-only or encoder-decoder? interpreting language model as a regularized encoder-decoder, 2023
Z. Fu, W. Lam, Q. Yu, A. M.-C. So, S. Hu, Z. Liu, and N. Collier · 2023
Closest in time.
Flax: A neural network library and ecosystem for JAX, 2023
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Protein language models trained on multiple sequence alignments learn phylogenetic relationships
U. Lupo, D. Sgarbossa, and A.-F. Bitbol · 2022
Cited alongside, same era.
Ensembl 2023
F. Martin et. al · 2022
Cited alongside, same era.
Progen2: Exploring the boundaries of protein language models
E. Nijkamp, J. A. Ruffolo, E. N. Weinstein, N. V. Naik, and A. Madani · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization, 2022
V. Sanh et. al · 2022
Cited alongside, same era.
Anvil - system architecture and experiences from deployment and early user operations
X. C. Song, P. Smith, R. Kalyanam, X. Zhu, E. Adams, K. Colby, P. Finnegan, E. Gough, E. Hillery, R. Irvine, A. Maji, and J. St. John · 2022
Cited alongside, same era.
Identifying and correcting repeat-calling errors in nanopore sequencing of telomeres
K.-T. Tan, M. K. Slevin, M. Meyerson, and H. Li · 2022
Cited alongside, same era.
ByT5: Towards a token-free future with pre-trained byte-to-byte models, Mar. 2022
L. Xue, A. Barua, N. Constant, R. Al-Rfou, S. Narang, M. Kale, A. Roberts, and C. Raffel · 2022
Cited alongside, same era.
Nucleotide transformer benchmark
InstaDeepAI · 2023
Closest in time.
TeloBase: a community-curated database of telomere sequences across the tree of life
M. Lyčka, M. Bubeník, M. Závodník, V. Peska, P. Fajkus, M. Demko, J. Fajkus, and M. Fojtová · 2023
Closest in time.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
E. D. Nguyen, M. Poli, M. Faizi, A. W. Thomas, C. J. Birch-sykes, M. Wornow, A. Patel, C. M. Rabideau, S. Massaroli, Y. Bengio, S. Ermon, S. A. Baccus, and C. Ré · 2023
Closest in time.
Foundation Models for Natural Language Processing
G. Paaß and S. Giesselbach · 2023
Closest in time.
DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks
D. Zhang, W. Zhang, B. He, J. Zhang, C. Qin, and J. Yao · 2023
Closest in time.
Dnabert-2: Efficient foundation model and benchmark for multi-species genome, 2023
Z. Zhou, Y. Ji, W. Li, P. Dutta, R. Davuluri, and H. Liu · 2023
Closest in time.
DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
Y. Ji, Z. Zhou, H. Liu, and R. V. Davuluri · 2059
Closest in time.
Transformer for Gene Expression Modeling (T-GEM): An Interpretable Deep Learning Model for Gene Expression-Based Phenotype Predictions
T.-H. Zhang, M. M. Hasib, Y.-C. Chiu, Z.-F. Han, Y.-F. Jin, M. Flores, Y. Chen, and Y. Huang · 2072
Closest in time.