Fetching the paper…
Reading the bibliography…
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 1907
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu · 2002
Earlier work this paper cites.
Understanding the difficulty of training transformers
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han · 2004
Earlier work this paper cites.
Adversarial training for large neural language models
X. Liu, H. Cheng, P. He, W. Chen, Y. Wang, H. Poon, and J. Gao · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
I. Dagan, O. Glickman, and B. Magnini · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
W. B. Dolan and C. Brockett · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
R. B. Haim, I. Dagan, B. Dolan, L. Ferro, D. Giampiccolo, B. Magnini, and I. Szpektor · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
D. Giampiccolo, B. Magnini, I. Dagan, and B. Dolan · 2007
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
L. Xiong, C. Xiong, Y. Li, K.-F. Tang, J. Liu, P. Bennett, J. Ahmed, and A. Overwijk · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs, 2017
I. Shankar, D. Nikhil, and C. Kornél · 2017
Earlier work this paper cites.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
T. Miyato, S.-i. Maeda, M. Koyama, and S. Ishii · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning
T. H. Trinh and Q. V. Le · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. Bowman · 2018
Earlier work this paper cites.
Taming sparsely activated transformer with stochastic experts
S. Zuo, X. Liu, J. Jiao, Y. J. Kim, H. Hassan, R. Zhang, T. Zhao, and J. Gao · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Openwebtext corpus
A. Gokaslan and V. Cohen · 2019
Cited alongside, same era.
H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts
M. Artetxe, S. Bhosale, N. Goyal, T. Mihaylov, M. Ott, S. Shleifer, X. V. Lin, J. Du, S. Iyer, R. Pasunuru, et al · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. v. d. Driessche, J.-B. Lespiau, B. Damoc, A. Clark, et al · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
L. Gao and J. Callan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in nlp
E. Strubell, A. Ganesh, and A. McCallum · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Cited alongside, same era.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. R. Bowman · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Cited alongside, same era.
L. Gui, B. Wang, Q. Huang, A. Hauptmann, Y. Bisk, and J. Gao · 2021
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
P. He, X. Liu, J. Gao, and W. Chen · 2021
Later among the works it cites.
How to train bert with an academic budget
P. Izsak, M. Berchansky, and O. Levy · 2021
Later among the works it cites.
COCO-LM: Correcting and contrasting text sequences for language model pretraining
Y. Meng, C. Xiong, P. Bajaj, S. Tiwary, P. Bennett, J. Han, and X. Song · 2021
Later among the works it cites.
Deep learning–based text classification: a comprehensive review
S. Minaee, N. Kalchbrenner, E. Cambria, N. Nikzad, M. Chenaghlu, and J. Gao · 2021
Later among the works it cites.
Do transformer modifications transfer across implementations and applications?
S. Narang, H. W. Chung, Y. Tay, W. Fedus, T. Fevry, M. Matena, K. Malkan, N. Fiedel, N. Shazeer, Z. Lan, et al · 2021
Later among the works it cites.
Large dual encoders are generalizable retrievers
J. Ni, C. Qu, J. Lu, Z. Dai, G. H. Ábrego, J. Ma, V. Y. Zhao, Y. Luan, K. B. Hall, M.-W. Chang, et al · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Later among the works it cites.
Training electra augmented with multi-word selection
J. Shen, J. Liu, T. Liu, C. Yu, and J. Han · 2021
Later among the works it cites.
Normformer: Improved transformer pretraining with extra normalization
S. Shleifer, J. Weston, and M. Ott · 2021
Later among the works it cites.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Y. Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y. Zhao, Y. Lu, et al · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
G. Yang, E. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao · 2021
Later among the works it cites.
Joint retrieval and generation training for grounded text generation
Y. Zhang, S. Sun, X. Gao, Y. Fang, C. Brockett, M. Galley, J. Gao, and B. Dolan · 2021
Later among the works it cites.
Neural approaches to conversational information retrieval
J. Gao, C. Xiong, P. Bennett, and N. Craswell · 2022
Closest in time.
C. Liang, H. Jiang, S. Zuo, P. He, X. Liu, J. Gao, W. Chen, and T. Zhao · 2022
Closest in time.
Pretraining text encoders with adversarial mixture of training signal generators
Y. Meng, C. Xiong, P. Bajaj, P. N. Bennett, J. Han, X. Song, et al · 2022
Closest in time.
Text and code embeddings by contrastive pre-training
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, et al · 2022
Closest in time.
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, et al · 2022
Closest in time.
Designing effective sparse expert models
B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus · 2022
Closest in time.