Fetching the paper…
Reading the bibliography…
Large Transformer models have been central to recent advances in natural language processing.
Evolutionary principles in self-referential learning. (on learning how to learn: The meta-meta-… hook.)
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
Evolving artificial neural networks
Xin Yao · 1999
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Zohar Karnin, Tomer Koren, and Oren Somekh · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, P. Barham, J. Chen, Z. Chen, Andy Davis, J. Dean, M. Devin, Sanjay Ghemawat, Geoffrey Irving, M. Isard, M. Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, D. Murray, Benoit Steiner, P. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Y. Yu, and Xiaoqiang Zhang · 2016
Earlier work this paper cites.
Dense associative memory for pattern recognition
Dmitry Krotov and John J. Hopfield · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Quoc V. Le, and Alex Kurakin · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Hierarchical representations for efficient architecture search
Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Tensor2tensor for neural machine translation
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N. Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit · 2018
Earlier work this paper cites.
QANet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le · 2018
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2018
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Program synthesis using uniform mutation by addition and deletion
Thomas Helmuth, N. McPhee, and L. Spector · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam M. Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and J. Richardson · 2018
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Evolving normalization-activation layers
Hanxiao Liu, Andrew Brock, Karen Simonyan, and Quoc V Le · 2020
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and T. Lillicrap · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The evolved transformer
David R. So, Chen Liang, and Quoc V. Le · 2019
Cited alongside, same era.
Designing neural networks through neuroevolution
Kenneth Stanley, Jeff Clune, Joel Lehman, and Risto Miikkulainen · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, R. Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, and Quoc V. Le · 2019
Cited alongside, same era.
Proxylessnas: Direct neural architecture search on target task and hardware
Han Cai, Ligeng Zhu, and Song Han · 2019
Cited alongside, same era.
Efficient multi-objective neural architecture search via lamarckian evolution
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Cited alongside, same era.
Later among the works it cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
Later among the works it cites.
Evaluating the search phase of neural architecture search
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat, and Mathieu Salzmann · 2020
Later among the works it cites.
Can weight sharing outperform random architecture search? An investigation with tunas
Gabriel Bender, Hanxiao Liu, Bo Chen, Grace Chu, Shuyang Cheng, Pieter-Jan Kindermans, and Quoc V. Le · 2020
Later among the works it cites.
Automl-zero: Evolving machine learning algorithms from scratch
Esteban Real, Chen Liang, David R. So, and Quoc V. Le · 2020
Later among the works it cites.
Multiplicative interactions and where to find them
Siddhant M. Jayakumar, Jacob Menick, Wojciech M. Czarnecki, Jonathan Schwarz, Jack W. Rae, Simon Osindero, Y. Teh, Tim Harley, and Razvan Pascanu · 2020
Later among the works it cites.
Glu variants improve transformer
Noam Shazeer · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, S. Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, L. Wang, and T. Liu · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam M. Shazeer · 2021
Closest in time.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2021
Closest in time.
Entailment as few-shot learner
Sinong Wang, Han Fang, Madian Khabsa, Hanzi Mao, and Hao Ma · 2021
Closest in time.
CvT: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Closest in time.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc V. Le, Chen Liang, Lluís-Miquel Munguía, D. Rothchild, David R. So, Maud Texier, and J. Dean · 2021
Closest in time.
https://www.gstatic.com/gumdrop/sustainability/24x7-carbon-free-energy-methodologies-metrics.pdf
24/7 carbon-free energy: Methodologies and metrics · 2021
Closest in time.