Fetching the paper…
Reading the bibliography…
Transformer-based large language models have displayed impressive in-context learning capabilities, where a pre-trained model can handle new tasks without fine-tuning by simply augmenting the query with some input-output examples from that task.
Concise formulas for the area and volume of a hyperspherical cap
Li, S · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Luo, J.-H., Wu, J., and Lin, W · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A · 2018
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
An improved analysis of training over-parameterized deep neural networks
Zou, D. and Gu, Q · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
A generalized neural tangent kernel analysis for two-layer neural networks
Chen, Z., Cao, Y., Gu, Q., and Zhang, T · 2020
Earlier work this paper cites.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Earlier work this paper cites.
Guaranteed recovery of one-hidden-layer neural networks via cross entropy
Fu, H., Chi, Y., and Liang, Y · 2020
Earlier work this paper cites.
An optimization and generalization analysis for max-pooling networks
Brutzkus, A. and Globerson, A · 2021
Earlier work this paper cites.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Karp, S., Winston, E., Li, Y., and Singh, A · 2021
Earlier work this paper cites.
Gender and representation bias in gpt-3 generated stories
Lucy, L. and Bamman, D · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Shi, Z., Wei, J., and Liang, Y · 2021
Cited alongside, same era.
Why lottery ticket wins? a theoretical perspective of sample complexity on sparse neural networks
Zhang, S., Wang, M., Liu, S., Chen, P.-Y., and Xiong, J · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
In-context convergence of transformers
Huang, Y., Cheng, Y., and Liang, Y · 2023
Later among the works it cites.
What improves the generalization of graph transformer? a theoretical dive into self-attention and positional encoding
Li, H., Wang, M., Ma, T., Liu, S., ZHANG, Z., and Chen, P.-Y · 2023
Later among the works it cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Later among the works it cites.
LLM-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
Cited alongside, same era.
Learning and generalization of one-hidden-layer neural networks, going beyond standard gaussian data
Li, H., Zhang, S., and Wang, M · 2022
Cited alongside, same era.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Liu, A., Swayamdipta, S., Smith, N. A., and Choi, Y · 2022
Cited alongside, same era.
What makes good in-context examples for gpt-3?
Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W · 2022
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Cited alongside, same era.
Oymak, S., Rawat, A. S., Soltanolkotabi, M., and Thrampoulidis, C · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
On the neural tangent kernel analysis of randomly pruned neural networks
Yang, H. and Wang, Z · 2023
Later among the works it cites.
Theoretical characterization of how neural network pruning affects its generalization
Yang, H., Liang, Y., Guo, X., Wu, L., and Wang, Z · 2023
Later among the works it cites.
Provably learning a multi-head attention layer
Chen, S. and Li, Y · 2024
Closest in time.
A provably effective method for pruning experts in fine-tuned sparse mixture-of-experts
Chowdhury, M. N. R., Wang, M., Maghraoui, K. E., Wang, N., Chen, P.-Y., and Carothers, C · 2024
Closest in time.
Superiority of multi-head attention in in-context linear regression
Cui, Y., Ren, J., He, P., Tang, J., and Xing, Y · 2024
Closest in time.
How does promoting the minority fraction affect generalization? a theoretical study of one-hidden-layer neural network on group imbalance
Li, H., Zhang, S., Zhang, Y., Wang, M., Liu, S., and Chen, P.-Y · 2024
Closest in time.
Enhancing graph transformers with hierarchical distance structural encoding, 2024
Luo, Y., Li, H., Shi, L., and Wu, X.-M · 2024
Closest in time.
How do skip connections affect graph convolutional networks with graph sampling? a theoretical analysis on generalization, 2024
Sun, J., Li, H., and Wang, M · 2024
Closest in time.
Do efficient transformers really save computation?
Yang, K., Ackermann, J., He, Z., Feng, G., Zhang, B., Feng, Y., Ye, Q., He, D., and Wang, L · 2024
Closest in time.
Visual prompting reimagined: The power of activation prompts, 2024
Zhang, Y., Li, H., Yao, Y., Chen, A., Zhang, S., Chen, P.-Y., Wang, M., and Liu, S · 2024
Closest in time.