Fetching the paper…
Reading the bibliography…
Large pretrained language models have achieved state-of-the-art results on a variety of downstream tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bram Bakker. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman Tijmen and Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Bidirectional lstm-crf models for sequence tagging
Zhiheng Huang, Wei Xu, and Kai Yu. 2015 · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis. 2017 · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le. 2017 · 2017
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Teacher guided architecture search
Pouya Bashivan, Mark Tensen, and James J DiCarlo. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
David R. So, Chen Liang, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Adabert: Task-adaptive bert compression with differentiable neural architecture search
Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. 2021 · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Later among the works it cites.
Annealing knowledge distillation
Aref Jafari, Mehdi Rezagholizadeh, Pranav Sharma, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
Improving task-agnostic bert distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2021 · 2021
Later among the works it cites.
Xtremedistiltransformers: Task transfer for task-agnostic distillation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
BERT-EMD: Many-to-many layer mapping for BERT compression with earth mover’s distance
Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020 · 2020
Cited alongside, same era.
Search to distill: Pearls are everywhere but not the eyes
Yu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli, Yukun Zhu, Bradley Green, and Xiaogang Wang. 2020 · 2020
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020 · 2020
Cited alongside, same era.
Subhabrata Mukherjee, Ahmed Hassan Awadallah, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
Accelerating neural architecture search via proxy data
Byunggook Na, Jisoo Mok, Hyeokjun Choe, and Sungroh Yoon. 2021 · 2021
Later among the works it cites.
Resources and benchmark corpora for hate speech detection: a systematic review
Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and Viviana Patti. 2021 · 2021
Later among the works it cites.
MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021 · 2021
Later among the works it cites.
Teacher guided neural architecture search for face recognition
Xiaobo Wang. 2021 · 2021
Later among the works it cites.
Universal-KD: Attention-based output-grounded intermediate layer knowledge distillation
Yimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md Akmal Haidar, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
Nas-bert: Task-agnostic and adaptive-size bert compression with neural architecture search
Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu. 2021 · 2021
Later among the works it cites.
Ibm announces new foundation model capabilities
Niklas Heidloff. 2023 · 2023
Closest in time.
Jongwoo Ko, Seungjoon Park, Minchan Jeong, Sukjin Hong, Euijai Ahn, Du-Seong Chang, and Se-Young Yun. 2023 · 2023
Closest in time.
Fair is fast, and fast is fair: Ibm slate foundation models for nlp
Alexander Lang. 2023 · 2023
Closest in time.
A comparative analysis of task-agnostic distillation methods for compressing transformer language models
Takuma Udagawa, Aashka Trivedi, Michele Merler, and Bishwaranjan Bhattacharjee. 2023 · 2023
Closest in time.
How to distill your BERT: An empirical study on the impact of weight initialisation and distillation objectives
Xinpeng Wang, Leonie Weissweiler, Hinrich Schütze, and Barbara Plank. 2023 · 2023
Closest in time.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. 2018 · 2023
Closest in time.