Fetching the paper…
Reading the bibliography…
Transformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
A new measure of rank correlation
Kendall, M. G. 1938 · 1938
Earlier work this paper cites.
A comparative analysis of selection schemes used in genetic algorithms
Goldberg, D.; and Deb, K. 1991 · 1991
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P.; Cesa-Bianchi, N.; and Fischer, P. 2002 · 2002
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C.; Ng, A.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y.; Schuster, M.; Chen, Z.; Le, Q.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. 2016 · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B.; and Le, Q. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017 · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y.; Fan, A.; Auli, M.; and Grangier, D. 2017 · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L.; Jamieson, K.; DeSalvo, G.; Rostamizadeh, A.; and Talwalkar, A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Quora question pairs
Chen, Z.; Zhang, H.; Zhang, X.; and Zhao, L. 2018 · 2018
Earlier work this paper cites.
DARTS: Differentiable Architecture Search
Liu, H.; Simonyan, K.; and Yang, Y. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Cited alongside, same era.
Efficient Neural Architecture Search via Parameter Sharing
Pham, H.; Guan, M.; Zoph, B.; Le, Q.; and Dean, J. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018 · 2018
The Evolved Transformer
So, D.; Le, Q.; and Liang, C. 2019 · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
Warstadt, A.; Singh, A.; and Bowman, S. 2019 · 2019
Later among the works it cites.
Auto-fpn: Automatic network architecture adaptation for object detection beyond classification
Xu, H.; Yao, L.; Zhang, W.; Liang, X.; and Li, Z. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; and Le, Q. 2019 · 2019
Later among the works it cites.
EENA: efficient evolution of neural architecture
Zhu, H.; An, Z.; Yang, C.; Xu, K.; Zhao, E.; and Xu, Y. 2019 · 2019
Later among the works it cites.
AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
Chen, D.; Li, Y.; Qiu, M.; Wang, Z.; Li, B.; Ding, B.; Deng, H.; Huang, J.; Lin, W.; and Zhou, J. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-Granularity Hierarchical Attention Fusion Networks for Reading Comprehension and Question Answering
Wang, W.; Yan, M.; and Wu, C. ???? · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
Pay Less Attention with Lightweight and Dynamic Convolutions
Wu, F.; Fan, A.; Baevski, A.; Dauphin, Y.; and Auli, M. 2018 · 2018
Cited alongside, same era.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
Zellers, R.; Bisk, Y.; Schwartz, R.; and Choi, Y. 2018 · 2018
Cited alongside, same era.
Once-for-All: Train One Network and Specialize it for Efficient Deployment
Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; and Han, S. 2019 · 2019
Cited alongside, same era.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K.; Luong, M.; Le, Q.; and Manning, C. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
DynaBERT: Dynamic BERT with Adaptive Width and Depth
Hou, L.; Huang, Z.; Shang, L.; Jiang, X.; Chen, X.; and Liu, Q. 2020 · 2020
Later among the works it cites.
ConvBERT: Improving BERT with Span-based Dynamic Convolution
Jiang, Z.; Yu, W.; Zhou, D.; Chen, Y.; Feng, J.; and Yan, S. 2020 · 2020
Later among the works it cites.
Reformer: The Efficient Transformer
Kitaev, N.; Kaiser, L.; and Levskaya, A. 2020 · 2020
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2020 · 2020
Later among the works it cites.
AutoML-zero: evolving machine learning algorithms from scratch
Real, E.; Liang, C.; So, D.; and Le, Q. 2020 · 2020
Later among the works it cites.
Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONAS
Shi, H.; Pi, R.; Xu, H.; Li, Z.; Kwok, J.; and Zhang, T. 2020 · 2020
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y.; Cordonnier, J.; and Loukas, A. 2021 · 2021
Closest in time.
Improve Vision Transformers Training by Suppressing Over-smoothing
Gong, C.; Wang, D.; Li, M.; Chandra, V.; and Liu, Q. 2021 · 2021
Closest in time.
Synthesizer: Rethinking self-attention in transformer models
Tay, Y.; Bahri, D.; Metzler, D.; Juan, D.; Zhao, Z.; and Zheng, C. 2021 · 2021
Closest in time.
Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
Yao, L.; Pi, R.; Xu, H.; Zhang, W.; Li, Z.; and Zhang, T. 2021 · 2021
Closest in time.