Fetching the paper…
Reading the bibliography…
Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications.
Statistical inference for probabilistic functions of finite state markov chains
Baum, L. E. and Petrie, T · 1966
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
Viterbi, A · 1967
Earlier work this paper cites.
Extensions of lipschitz mappings into a hilbert space
Johnson, W. B. and Lindenstrauss, J · 1984
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Approximate nearest neighbor queries in fixed dimensions
Arya, S. and Mount, D. M · 1993
Earlier work this paper cites.
The space complexity of approximating the frequency moments
Alon, N., Matias, Y., and Szegedy, M · 1996
Earlier work this paper cites.
A study of branch prediction strategies
Smith, J. E · 1998
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Gionis, A., Indyk, P., Motwani, R., et al · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Laurent, B. and Massart, P · 2000
Earlier work this paper cites.
Finding frequent items in data streams
Charikar, M., Chen, K., and Farach-Colton, M · 2002
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C · 2003
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
Datar, M., Immorlica, N., Indyk, P., and Mirrokni, V. S · 2004
Earlier work this paper cites.
Mean shift clustering
Derpanis, K. G · 2005
Earlier work this paper cites.
Improved approximation algorithms for large matrices via random projections
Sarlos, T · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
Multidimensional scaling, 315–347
Cox, M. and Cox, T · 2008
Earlier work this paper cites.
Subspace embeddings for the l1-norm with applications
Sohler, C. and Woodruff, D. P · 2011
Earlier work this paper cites.
CUDA Programming: A Developer’s Guide to Parallel Computing with GPUs
Cook, S · 2012
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Gordon, A., Kozareva, Z., and Roemmele, M · 2012
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Clarkson, K. L. and Woodruff, D. P · 2013
Earlier work this paper cites.
How to access global memory efficiently in CUDA C/C++ kernels
Harris, M · 2013
Earlier work this paper cites.
Faster ridge regression via the subsampled randomized hadamard transform
Lu, Y., Dhillon, P., Foster, D. P., and Ungar, L · 2013
Earlier work this paper cites.
Low-distortion subspace embeddings in input-sparsity time and applications to robust linear regression
Meng, X. and Mahoney, M. W · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Nelson, J. and Nguyên, H. L · 2013
Earlier work this paper cites.
Beyond locality-sensitive hashing
Andoni, A., Indyk, P., Nguyen, H. L., and Razenshteyn, I · 2014
Earlier work this paper cites.
Approximate nearest neighbor algorithm based on navigable small world graphs
Malkov, Y., Ponomarenko, A., Logvinov, A., and Krylov, V · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
Woodruff, D. P · 2014
Earlier work this paper cites.
Optimal data-dependent hashing for approximate near neighbors
Andoni, A. and Razenshteyn, I · 2015
Earlier work this paper cites.
Practical and optimal lsh for angular distance
Andoni, A., Indyk, P., Laarhoven, T., Razenshteyn, I., and Schmidt, L · 2015
Earlier work this paper cites.
Fast and accurate maximum inner product recommendations on map-reduce
Hall, R. and Attenberg, J · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
On symmetric and asymmetric lshs for inner product search
Neyshabur, B. and Srebro, N · 2015
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Boutsidis, C., Woodruff, D. P., and Zhong, P · 2016
Earlier work this paper cites.
Off the beaten path: Let’s replace term-based retrieval with k-nn search
Boytsov, L., Novak, D., Malkov, Y., and Nyberg, E · 2016
Earlier work this paper cites.
Nearly tight oblivious subspace embeddings by trace inequalities
Cohen, M. B · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Reasoning in vector space: An exploratory study of question answering
Lee, M., He, X., Yih, W.-t., Gao, J., Deng, L., and Smolensky, P · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Razenshteyn, I., Song, Z., and Woodruff, D. P · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Veit, A., Wilber, M. J., and Belongie, S · 2016
Earlier work this paper cites.
Optimal hashing-based time-space trade-offs for approximate near neighbors
Andoni, A., Laarhoven, T., Razenshteyn, I., and Waingarten, E · 2017
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B · 2017
Cited alongside, same era.
Low rank approximation with entrywise l1-norm error
Song, Z., Woodruff, D. P., and Zhong, P · 2017
Cited alongside, same era.
Deep matrix factorization models for recommender systems
Xue, H.-J., Dai, X., Zhang, J., Huang, S., and Chen, J · 2017
Cited alongside, same era.
Approximate nearest neighbor search in high dimensions
Andoni, A., Indyk, P., and Razenshteyn, I · 2018
Cited alongside, same era.
On the hardness of approximate and exact (bichromatic) maximum inner product
Chen, L · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Turbotransformers: an efficient gpu serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Later among the works it cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Later among the works it cites.
The hardware lottery
Hooker, S · 2021
Later among the works it cites.
Data movement is all you need: A case study on optimizing transformers
Ivanov, A., Dryden, N., Ben-Nun, T., Li, S., and Hoefler, T · 2021
Later among the works it cites.
A faster algorithm for solving general lps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Approximate nearest neighbors in limited space
Indyk, P. and Wagner, T · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Cited alongside, same era.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Cited alongside, same era.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Malkov, Y. A. and Yashunin, D. A · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Cited alongside, same era.
Jiang, S., Song, Z., Weinstein, O., and Zhang, H · 2021
Later among the works it cites.
Sublinear least-squares value iteration via locality sensitive hashing
Shrivastava, A., Song, Z., and Xu, Z · 2021
Later among the works it cites.
Oblivious sketching-based central path method for linear programming
Song, Z. and Yu, Z · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
Song, Z., Zhang, L., and Zhang, R · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
GPT-J-6B: A 6 billion parameter autoregressive language model
Wang, B. and Komatsuzaki, A · 2021
Later among the works it cites.
Lightseq: A high performance inference library for transformers
Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L · 2021
Later among the works it cites.
Alman, J., Liang, J., Song, Z., Zhang, R., and Zhuo, D · 2022
Later among the works it cites.
Deepspeed-inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al · 2022
Later among the works it cites.
Bansal, H., Gopalakrishnan, K., Dingliwal, S., Bodapati, S., Kirchhoff, K., and Roth, D · 2022
Later among the works it cites.
GPT-NeoX-20B: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Chan, S. C., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A. K., Richemond, P. H., McClelland, J., and Hill, F · 2022
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Gu, Y. and Song, Z · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Hu, H., Song, Z., Weinstein, O., and Zhuo, D · 2022
Later among the works it cites.
C-MinHash: Improving minwise hashing with circulant permutation
Li, X. and Li, P · 2022
Later among the works it cites.
Large models are parsimonious learners: Activation sparsity in trained transformers, 2022
Li, Z., You, C., Bhojanapalli, S., Li, D., Rawat, A. S., Reddi, S. J., Ye, K., Chern, F., Yu, F., Guo, R., and Kumar, S · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
Halos: Hashing large output space for cheap inference
Liu, Z., Xu, Z., Ji, A., Zhang, J., Li, J., Chen, B., and Shrivastava, A · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Later among the works it cites.
Gpu performance background user’s guide, 2022
NVIDIA · 2022
Later among the works it cites.
nuqmm: Quantized matmul for efficient inference of large-scale generative language models
Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D · 2022
Later among the works it cites.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Zhang, L · 2022
Later among the works it cites.
Fast attention requires bounded entries
Alman, J. and Song, Z · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Brand, J. v. d., Song, Z., and Zhou, T · 2023
Closest in time.
Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Gu, Y., Song, Z., Yin, J., and Zhang, L · 2023
Closest in time.
Efficient asynchronize stochastic gradient algorithm with structured data
Song, Z. and Ye, M · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Zandieh, A., Han, I., Daliri, M., and Karbasi, A · 2023
Closest in time.