Fetching the paper…
Reading the bibliography…
Recently, large pre-trained models have significantly improved the performance of various Natural LanguageProcessing (NLP) tasks but they are expensive to serve due to long serving latency and large memory usage.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M · 2009
Earlier work this paper cites.
Firefly algorithm, stochastic test functions and design optimisation
Yang, X.-S · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H. H., and Leyton-Brown, K · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and De Freitas, N · 2015
Earlier work this paper cites.
Scalable bayesian optimization using deep neural networks
Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M., Prabhat, M., and Adams, R · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Bayesian optimization for automated model selection
Malkomes, G., Schaff, C., and Garnett, R · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Deep kernel learning
Wilson, A. G., Hu, Z., Salakhutdinov, R., and Xing, E. P · 2016
Earlier work this paper cites.
Boat: Building auto-tuners with structured bayesian optimization
Dalibard, V., Schaarschmidt, M., and Yoneki, E · 2017
Earlier work this paper cites.
Google vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D · 2017
Earlier work this paper cites.
Learning to select data for transfer learning with bayesian optimization
Ruder, S. and Plank, B · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Bayesian optimization and data science
Archetti, F. and Candelieri, A · 2019
Cited alongside, same era.
Ernie: Enhanced language representation with informative entities
Zhang, Z., Han, X., Liu, Z., Jiang, X., Sun, M., and Liu, Q · 2019
Later among the works it cites.
AdaBERT: Task-adaptive BERT compression with differentiable neural architecture search
Chen, D., Li, Y., Qiu, M., Wang, Z., Li, B., Ding, B., Deng, H., Huang, J., Lin, W., and Zhou, J · 2020
Later among the works it cites.
Gpt-3: Its nature, scope, limits, and consequences
Floridi, L. and Chiriatti, M · 2020
Later among the works it cites.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Gordon, M., Duh, K., and Andrews, N · 2020
Later among the works it cites.
DynaBERT: Dynamic bert with adaptive width and depth
Hou, L., Huang, Z., Shang, L., Jiang, X., Chen, X., and Liu, Q · 2020
Later among the works it cites.
TinyBERT: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Software-defined far memory in warehouse-scale computers
Lagar-Cavilla, A., Ahn, J., Souhlal, S., Agarwal, N., Burny, R., Butt, S., Chang, J., Chaugule, A., Deng, N., Shahid, J., et al · 2019
Cited alongside, same era.
Mlperf training benchmark, 2019
Mattson, P., Cheng, C., Coleman, C., Diamos, G., Micikevicius, P., Patterson, D., Tang, H., Wei, G.-Y., Bailis, P., Bittorf, V., Brooks, D., Chen, D., Dutta, D., Gupta, U., Hazelwood, K., Hock, A., Huang, X., Ike, A., Jia, B., Kang, D., Kanter, D., Kumar, N., Liao, J., Ma, G., Narayanan, D., Oguntebi, T., Pekhimenko, G., Pentecost, L., Reddi, V. J., Robie, T., John, T. S., Tabaru, T., Wu, C.-J., Xu, L., Yamazaki, M., Young, C., and Zaharia, M · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Distilling task-specific knowledge from BERT into simple neural networks
Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J · 2019
Cited alongside, same era.
Small and practical BERT models for sequence labeling
Tsai, H., Riesa, J., Johnson, M., Arivazhagan, N., Li, X., and Archer, A · 2019
Cited alongside, same era.
Later among the works it cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Later among the works it cites.
Q-BERT: Hessian based ultra low precision quantization of BERT
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Later among the works it cites.
Finding fast transformers: One-shot neural architecture search by component composition
Tsai, H., Ooi, J., Ferng, C.-S., Chung, H. W., and Riesa, J · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Later among the works it cites.
AutoBERT-Zero: Evolving bert backbone from scratch
Gao, J., Xu, H., Ren, X., Yu, P. L., Liang, X., Jiang, X., Li, Z., et al · 2021
Later among the works it cites.
Ten lessons from three generations shaped google’s tpuv4i: Industrial product
Jouppi, N. P., Yoon, D. H., Ashcraft, M., Gottscho, M., Jablin, T. B., Kurian, G., Laudon, J., Li, S., Ma, P., Ma, X., et al · 2021
Later among the works it cites.
Searching for fast model families on datacenter accelerators
Li, S., Tan, M., Pang, R., Li, A., Cheng, L., Le, Q. V., and Jouppi, N. P · 2021
Later among the works it cites.
NAS-BERT: Task-agnostic and adaptive-size bert compression with neural architecture search
Xu, J., Tan, X., Luo, R., Song, K., Li, J., Qin, T., and Liu, T.-Y · 2021
Later among the works it cites.