Fetching the paper…
Reading the bibliography…
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP.
Piqa: An algebra for querying protein data sets
S. Tata and J. M. Patel · 2003
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Recognizing textual entailment: Models and applications
I. Dagan, D. Roth, M. Sammons, and F. M. Zanzotto · 2013
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs, 2017
S. Iyer, N. Dandekar, and K. Csernai · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. R. Bowman · 2017
Earlier work this paper cites.
Copa: Constrained parafac2 for sparse & large datasets
A. Afshar, I. Perros, E. E. Papalexakis, E. Searles, J. Ho, and J. Sun · 2018
Earlier work this paper cites.
A systematic classification of knowledge, reasoning, and context within the arc dataset
M. Boratko, H. Padigela, D. Mikkilineni, P. Yuvraj, R. Das, A. McCallum, M. Chang, A. Fokoue-Nkoutche, P. Kapanipathi, N. Mattei, et al · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Cited alongside, same era.
Efficient training of bert by progressively stacking
L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2019
Power-bert: Accelerating bert inference via progressive word-vector elimination
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma · 2020
Later among the works it cites.
Length-adaptive transformer: Train once with length drop, use anytime with search
G. Kim and K. Cho · 2020
Later among the works it cites.
Shallow-to-deep training for neural machine translation
B. Li, Z. Wang, H. Liu, Y. Jiang, Q. Du, T. Xiao, H. Wang, and J. Zhu · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
J. Vig and Y. Belinkov · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov · 2019
Cited alongside, same era.
Pytorch image models
R. Wightman · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi · 2020
Later among the works it cites.
Anlizing the adversarial natural language inference dataset
A. Williams, T. Thrush, and D. Kiela · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
M. Zhang and Y. He · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Ast: Audio spectrogram transformer
Y. Gong, Y.-A. Chung, and J. Glass · 2021
Later among the works it cites.
Pct: Point cloud transformer
M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu · 2021
Later among the works it cites.
Learned token pruning for transformers
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer · 2021
Later among the works it cites.
C. Li, M. Zhang, and Y. He · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith, and M. Lewis · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
H. Wang, Z. Zhang, and S. Han · 2021
Later among the works it cites.
Token dropping for efficient BERT pretraining
L. Hou, R. Y. Pang, T. Zhou, Y. Wu, X. Song, X. Song, and D. Zhou · 2022
Closest in time.
Staged training for transformer language models
S. Shen, P. Walsh, K. Keutzer, J. Dodge, M. Peters, and I. Beltagy · 2022
Closest in time.