Fetching the paper…
Reading the bibliography…
Foundation models are applied in a broad spectrum of settings with different inference constraints, from massive multi-accelerator clusters to resource-constrained standalone mobile devices.
Universally slimmable networks and improved training techniques, 2019a
J. Yu and T. Huang · 1903
Earlier work this paper cites.
Hat: Hardware-aware transformers for efficient natural language processing
H. Wang, Z. Wu, Z. Liu, H. Cai, L. Zhu, C. Gan, and S. Han · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The winograd schema challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
J. Berant, A. K. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context, 2016
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations, 2017
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning, 2017
B. Zoph and Q. V. Le · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering, 2018
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
N. Shazeer and M. Stern · 2018
Earlier work this paper cites.
J. Yu, L. Yang, N. Xu, J. Yang, and T. Huang · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2019
Earlier work this paper cites.
Once-for-all: Train one network and specialize it for efficient deployment
H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
Regularized evolution for image classifier architecture search, 2019
E. Real, A. Aggarwal, Y. Huang, and Q. V. Le · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?, 2019
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Dynamic convnets on tiny devices via nested sparsity
M. Grimaldi, L. Mocerino, A. Cipolletta, and A. Calimera · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Matryoshka representation learning
A. Kusupati, G. Bhatt, A. Rege, M. Wallingford, A. Sinha, V. Ramanujan, W. Howard-Snyder, K. Chen, S. Kakade, P. Jain, et al · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
M. Li, S. Gururangan, T. Dettmers, M. Lewis, T. Althoff, N. A. Smith, and L. Zettlemoyer · 2022
Later among the works it cites.
Confident adaptive language modeling
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Tran, Y. Tay, and D. Metzler · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Characterising bias in compressed models
S. Hooker, N. Moorosi, G. Clark, S. Bengio, and E. Denton · 2020
Cited alongside, same era.
Dynabert: Dynamic bert with adaptive width and depth
L. Hou, Z. Huang, L. Shang, X. Jiang, X. Chen, and Q. Liu · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Soft threshold weight reparameterization for learnable sparsity
A. Kusupati, V. Ramanujan, R. Somani, M. Wortsman, P. Jain, S. Kakade, and A. Farhadi · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding, 2020
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2020
Cited alongside, same era.
The low-resource double bind: An empirical study of pruning for low-resource machine translation
O. Ahia, J. Kreutzer, and S. Hooker · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Later among the works it cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Closest in time.
Flexivit: One model for all patch sizes
L. Beyer, P. Izmailov, A. Kolesnikov, M. Caron, S. Kornblith, X. Zhai, M. Minderer, M. Tschannen, I. Alabdulmohsin, and F. Pavetic · 2023
Closest in time.
Accelerating large language model decoding with speculative sampling
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Closest in time.
The case for 4-bit precision: k-bit inference scaling laws
T. Dettmers and L. Zettlemoyer · 2023
Closest in time.
Full stack optimization of transformer inference: a survey
S. Kim, C. Hooper, T. Wattanawong, M. Kang, R. Yan, H. Genc, G. Dinh, Q. Huang, K. Keutzer, M. W. Mahoney, et al · 2023
Closest in time.
Fast inference from transformers via speculative decoding
Y. Leviathan, M. Kalman, and Y. Matias · 2023
Closest in time.
Gpt-4 technical report
R. OpenAI · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Closest in time.
Sharcs: Efficient transformers through routing with dynamic width sub-networks
M. Salehi, S. Mehta, A. Kusupati, A. Farhadi, and H. Hajishirzi · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
M. Valipour, M. Rezagholizadeh, H. Rajabzadeh, M. Tahaei, B. Chen, and A. Ghodsi · 2023
Closest in time.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2023
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Hybrid llm: Cost-efficient and quality-aware query routing
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. Ruhle, L. V. Lakshmanan, and A. H. Awadallah · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.