Parallel prefix computation
R. E. Ladner and M. J. Fischer · 1980
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Introduction to Algorithms, 3rd Edition
T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein · 2009
Earlier work this paper cites.
Identification of direct residue contacts in protein–protein interaction by message passing
M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, and T. Hwa · 2009
Earlier work this paper cites.
Three-dimensional structures of membrane proteins from genomic sequencing
T. A. Hopf, L. J. Colwell, R. Sheridan, B. Rost, C. Sander, and D. S. Marks · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2014
Earlier work this paper cites.
Robust and accurate prediction of residue–residue interactions across protein interfaces using evolutionary information
S. Ovchinnikov, H. Kamisetty, and D. Baker · 2014
Earlier work this paper cites.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Earlier work this paper cites.
Inferring interaction partners from protein sequences
A.-F. Bitbol, R. S. Dwyer, L. J. Colwell, and N. S. Wingreen · 2016
Earlier work this paper cites.
Hierarchical attention networks for document classification
Z. Yang, D. Yang, C. Dyer, X. He, A. J. Smola, and E. H. Hovy · 2016
Earlier work this paper cites.
Orthogonal random features
F. X. Yu, A. T. Suresh, K. M. Choromanski, D. N. Holtmann-Rice, and S. Kumar · 2016
Earlier work this paper cites.
The unreasonable effectiveness of structured random orthogonal embeddings
K. M. Choromanski, M. Rowland, and A. Weller · 2017
Earlier work this paper cites.
Improved training of Wasserstein GANs
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. F. Foster, L. Jones, M. Schuster, N. Shazeer, N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, Z. Chen, Y. Wu, and M. Hughes · 2018
Earlier work this paper cites.
Initialization matters: Orthogonal predictive state recurrent neural networks
K. Choromanski, C. Downey, and B. Boots · 2018
Earlier work this paper cites.
The geometry of random features
K. Choromanski, M. Rowland, T. Sarlós, V. Sindhwani, R. E. Turner, and A. Weller · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Original
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
R. Frostig, M. Johnson, and C. Leary · 2018
Earlier work this paper cites.