Fetching the paper…
Reading the bibliography…
Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis.
The human genome browser at ucsc
W. J. Kent, C. W. Sugnet, T. S. Furey, K. M. Roskin, T. H. Pringle, A. M. Zahler, and D. Haussler · 2002
Earlier work this paper cites.
Qualitatively predicting acetylation and methylation areas in DNA sequences
T. H. Pham, D. H. Tran, T. B. H. Ho, K. Satou, and G. Valiente · 2005
Earlier work this paper cites.
Genome-wide map of nucleosome acetylation and methylation in yeast
D. K. Pokholok, C. T. Harbison, S. Levine, F. Lewitter, D. K. Gifford, and R. A. Young · 2005
Earlier work this paper cites.
Visualizing data using t-sne
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Genetics talks to epigenetics? the interplay between sequence variants and chromatin structure
S. Zaina, E. L. Pérez-Luque, and G. Lund · 2010
Earlier work this paper cites.
Modernizing reference genome assemblies
D. M. Church, V. A. Schneider, T. Graves, K. Auger, F. Cunningham, N. Bouk, H.-C. Chen, R. Agarwala, W. M. McLaren, G. R. Ritchie, et al · 2011
Earlier work this paper cites.
An integrated encyclopedia of dna elements in the human genome
ENCODE Project Consortium · 2012
Earlier work this paper cites.
Genome reference consortium human build 38 (grch38)
Genome Reference Consortium · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Comparative primate genomics: emerging patterns of genome content and dynamics
J. Rogers and R. A. Gibbs · 2014
Earlier work this paper cites.
Integrative analysis of 111 reference human epigenomes
Roadmap Epigenomics Consortium · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
J. Zhou and O. G. Troyanskaya · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
DeePromoter: Robust promoter predictor using deep learning
M. Oubounyt, Z. Louadi, H. Tayara, and K. T. Chong · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Expanded encyclopaedias of DNA elements in the human and mouse genomes
ENCODE Project Consortium · 2020
Earlier work this paper cites.
Shortformer: Better language modeling using shorter inputs
O. Press, N. A. Smith, and M. Lewis · 2020
Cited alongside, same era.
Transformer protein language models are unsupervised structure learners
R. Rao, J. Meier, T. Sercu, S. Ovchinnikov, and A. Rives · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Cited alongside, same era.
Effective gene expression prediction from sequence by integrating long-range interactions
Ž. Avsec, V. Agarwal, D. Visentin, J. R. Ledsam, A. Grabska-Barwinska, K. R. Taylor, Y. Assael, J. Jumper, P. Kohli, and D. R. Kelley · 2021
DNA language models are powerful zero-shot predictors of non-coding variant effects
G. Benegas, S. S. Batra, and Y. S. Song · 2022
Later among the works it cites.
ProteinBERT: a universal deep-learning model of protein sequence and function
N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, and M. Linial · 2022
Later among the works it cites.
Ensembl 2022
F. Cunningham, J. E. Allen, J. Allen, J. Alvarez-Jarreta, M. R. Amode, I. M. Armean, O. Austine-Orimoloye, A. G. Azov, I. Barnes, R. Bennett, et al · 2022
Later among the works it cites.
ProtGPT2 is a deep unsupervised language model for protein design
N. Ferruz, S. Schmidt, and B. Höcker · 2022
Later among the works it cites.
A deep learning framework for enhancer prediction using word embedding and sequence generation
Q. Geng, R. Yang, and L. Zhang · 2022
Later among the works it cites.
Genomic Benchmarks: A collection of datasets for genomic sequence classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Prottrans: Toward understanding the language of life through self-supervised learning
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, et al · 2021
Cited alongside, same era.
A practical survey on faster and lighter transformers
Q. Fournier, G. M. Caron, and D. Aloise · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2021
Cited alongside, same era.
DNABERT: pre-trained bidirectional encoder representations from transformers model for DNA-language in genome
Y. Ji, Z. Zhou, H. Liu, and R. V. Davuluri · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Cited alongside, same era.
Language models enable zero-shot prediction of the effects of mutations on protein function
J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, and A. Rives · 2021
Cited alongside, same era.
K. Gresova, V. Martinek, D. Cechak, P. Simecek, and P. Alexiou · 2022
Later among the works it cites.
The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models
C. Li, M. Zhang, and Y. He · 2022
Later among the works it cites.
Language models of protein sequences at the scale of evolution enable accurate structure prediction
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido, et al · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. W. Linderman · 2022
Later among the works it cites.
scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data
F. Yang, W. Wang, F. Wang, Y. Fang, D. Tang, J. Huang, H. Lu, and J. Yao · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics
M. Zvyagin, A. Brace, K. Hippe, Y. Deng, B. Zhang, C. O. Bohorquez, A. Clyde, B. Kale, D. Perez-Rivera, H. Ma, et al · 2022
Later among the works it cites.
The Nucleotide Transformer: Building and evaluating robust foundation models for human genomics
H. Dalla-Torre, L. Gonzalez, J. Mendoza-Revilla, N. L. Carranza, A. H. Grzywaczewski, F. Oteri, C. Dallago, E. Trop, H. Sirelkhatim, G. Richard, M. Skwark, K. Beguir, M. Lopez, and T. Pierrot · 2023
Closest in time.
Simple hardware-efficient long convolutions for sequence modeling
D. Y. Fu, E. L. Epstein, E. Nguyen, A. W. Thomas, M. Zhang, T. Dao, A. Rudra, and C. Ré · 2023
Closest in time.
Species-aware DNA language modeling
D. Gankin, A. Karollus, M. Grosshauser, K. Klemon, J. Hingerl, and J. Gagneur · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig · 2023
Closest in time.
Large language models generate functional protein sequences across diverse families
A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong, Z. Z. Sun, R. Socher, et al · 2023
Closest in time.
Hyena Hierarchy: Towards larger convolutional language models
M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. Ré · 2023
Closest in time.