Fetching the paper…
Reading the bibliography…
Pretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelligence.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Training binary neural networks with real-to-binary convolutions
Martinez, B.; Yang, J.; Bulat, A.; and Tzimiropoulos, G. 2020 · 2003
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2004
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Courbariaux, M.; Hubara, I.; Soudry, D.; El-Yaniv, R.; and Bengio, Y. 2016 · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Liu, Z.; Wu, B.; Luo, W.; Yang, X.; Liu, W.; and Cheng, K.-T. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Earlier work this paper cites.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning
Gordon, M.; Duh, K.; and Andrews, N. 2020 · 2020
Earlier work this paper cites.
ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
Liu, Z.; Shen, Z.; Savvides, M.; and Cheng, K.-T. 2020 · 2020
Earlier work this paper cites.
Cp-nas: Child-parent neural architecture search for 1-bit cnns
Li’an Zhuo, B. Z.; Chen, H.; Yang, L.; Chen, C.; Zhu, Y.; and Doermann, D. 2020 · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
BinaryBERT: Pushing the Limit of BERT Quantization
Bai, H.; Zhang, W.; Hou, L.; Shang, L.; Jin, J.; Jiang, X.; Liu, Q.; Lyu, M.; and King, I. 2021 · 2021
Cited alongside, same era.
I-bert: Integer-only bert quantization
Kim, S.; Gholami, A.; Yao, Z.; Mahoney, M. W.; and Keutzer, K. 2021 · 2021
Cited alongside, same era.
OPTQ: Accurate quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
Dq-bart: Efficient sequence-to-sequence model via joint distillation and quantization
Li, Z.; Wang, Z.; Tan, M.; Nallapati, R.; Bhatia, P.; Arnold, A.; Xiang, B.; and Roth, D. 2022 · 2022
Cited alongside, same era.
Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation
He, B.; Martens, J.; Zhang, G.; Botev, A.; Brock, A.; Smith, S. L.; and Teh, Y. W. 2023 · 2023
Closest in time.
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
He, P.; Gao, J.; and Chen, W. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Closest in time.
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping
Kunes, R. Z.; Yin, M.; Land, M.; Haviv, D.; Pe’er, D.; and Tavaré, S. 2023 · 2023
Closest in time.
Binary and Ternary Natural Language Generation
Liu, Z.; Oguz, B.; Pappu, A.; Shi, Y.; and Krishnamoorthi, R. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, M.; Ji, R.; Xu, Z.; Zhang, B.; Chao, F.; Lin, C.-W.; and Shao, L. 2022 · 2022
Cited alongside, same era.
Bit: Robustly binarized multi-distilled transformer
Liu, Z.; Oguz, B.; Pappu, A.; Xiao, L.; Yih, S.; Li, M.; Krishnamoorthi, R.; and Mehdad, Y. 2022 · 2022
Cited alongside, same era.
Bibert: Accurate fully binarized bert
Qin, H.; Ding, Y.; Zhang, M.; Yan, Q.; Liu, A.; Dang, Q.; Liu, Z.; and Liu, X. 2022 · 2022
Cited alongside, same era.
Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and Filtering
Xing, X.; Jiang, Y.; Zhang, B.; Ding, W.; Li, Y.; Li, H.; and Peng, H. 2022a · 2022
Cited alongside, same era.
Quantifying Memorization Across Neural Language Models
Carlini, N.; Ippolito, D.; Jagielski, M.; Lee, K.; Tramer, F.; and Zhang, C. 2023 · 2023
Cited alongside, same era.
An Equivalence Analysis of Binary Quantification Methods
Castano, A.; Alonso, J.; González, P.; and del Coz, J. J. 2023 · 2023
Cited alongside, same era.
CycleMLP: A MLP-like Architecture for Dense Visual Predictions
Chen, S.; Xie, E.; Ge, C.; Chen, R.; Liang, D.; and Luo, P. 2023 · 2023
Cited alongside, same era.
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Seggpt: Segmenting everything in context
Wang, X.; Zhang, X.; Cao, Y.; Wang, W.; Shen, C.; and Huang, T. 2023 · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Closest in time.
Xing, X.; Du, L.; Wang, X.; Zeng, X.; Wang, Y.; Zhang, Z.; and Zhang, J. 2023 · 2023
Closest in time.
Resilient Binary Neural Network
Xu, S.; Li, Y.; Ma, T.; Lin, M.; Dong, H.; Zhang, B.; Gao, P.; and Lu, J. 2023 · 2023
Closest in time.
Holistic Adversarially Robust Pruning
Zhao, Q.; and Wressnegger, C. 2023 · 2023
Closest in time.