Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have brought significant changes to human society.
The volumetric barrier for semidefinite programming
Kurt M Anstreicher · 2000
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Residual conv-deconv grid network for semantic segmentation
Damien Fourure, Rémi Emonet, Elisa Fromont, Damien Muselet, Alain Tremeau, and Christian Wolf · 2017
Earlier work this paper cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Reversible architectures for arbitrarily deep residual neural networks
Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, and Elliot Holtham · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Yiping Lu, Aoxiao Zhong, Quanzheng Li, and Bin Dong · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Classification of breast cancer histology images using transfer learning
Sulaiman Vesal, Nishant Ravikumar, AmirAbbas Davari, Stephan Ellmann, and Andreas Maier · 2018
Earlier work this paper cites.
Understanding convolution for semantic segmentation
Panqu Wang, Pengfei Chen, Ye Yuan, Ding Liu, Zehua Huang, Xiaodi Hou, and Garrison Cottrell · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Earlier work this paper cites.
Optimal sketching for kronecker product regression and low rank approximation
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Earlier work this paper cites.
Object detection based on vgg with resnet network
Md Foysal Haque, Hye-Youn Lim, and Dae-Seong Kang · 2019
Earlier work this paper cites.
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Object detection based on ssd-resnet
Xin Lu, Xin Kang, Shun Nishide, and Fuji Ren · 2019
Earlier work this paper cites.
Multi-object detection method based on yolo and resnet hybrid networks
Zhenyu Lu, Jia Lu, Quanbo Ge, and Tianming Zhan · 2019
Earlier work this paper cites.
Categorical image classification based on representational deep network (resnet)
Arpana Mahajan and Sanjay Chaudhary · 2019
Earlier work this paper cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric Villemonte de La Clergerie, Djamé Seddah, and Benoit Sagot · 2019
Earlier work this paper cites.
Moving object detection method via resnet-18 with encoder–decoder structure in complex scenes
Xianfeng Ou, Pengcheng Yan, Yiming Zhang, Bing Tu, Guoyun Zhang, Jianhui Wu, and Wujing Li · 2019
Earlier work this paper cites.
Dn-resnet: Efficient deep residual network for image denoising
Haoyu Ren, Mostafa El-Khamy, and Jungwon Lee · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Deep semantic segmentation of kidney and space-occupying lesion area based on scnn and resnet models combined with sift-flow algorithm
Kai-jian Xia, Hong-sheng Yin, and Yu-dong Zhang · 2019
Cited alongside, same era.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Cited alongside, same era.
Automatic traffic accident detection system using resnet and svm
Aman Kumar Agrawal, Kadamb Agarwal, Jitendra Choudhary, Aradhita Bhattacharya, Srihitha Tangudu, Nishkarsh Makhija, and B Rajitha · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Adore: Differentially oblivious relational database operators
Lianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song, Danyang Zhuo, and Shumo Chu · 2022
Later among the works it cites.
Adaptive and dynamic multi-resolution hashing for pairwise summations
Lianke Qin, Aravind Reddy, Zhao Song, Zhaozhuo Xu, and Danyang Zhuo · 2022
Later among the works it cites.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Research on animated gifs emotion recognition based on resnet-convgru
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2020
Cited alongside, same era.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Smyrf-efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis · 2020
Cited alongside, same era.
A faster interior point method for semidefinite programming
Haotian Jiang, Tarun Kathuria, Yin Tat Lee, Swati Padmanabhan, and Zhao Song · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Efficient resnet model to predict protein-protein interactions with gpu computing
Shengyu Lu, Qingqi Hong, Beizhan Wang, and Hongji Wang · 2020
Cited alongside, same era.
Qian Zhang, Ren Qing-Dao-Er-Ji, Na Li, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Faster robust tensor power method for arbitrary order
Yichuan Deng, Zhao Song, and Junze Yin · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Enhancing breast cancer classification using transfer resnet with lightweight attention mechanism
Suxing Liu · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Space-efficient interior point method, with applications to linear programming and maximum weight bipartite matching
S. Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2023
Closest in time.
A deep analysis of transfer learning based breast cancer detection using histopathology images
Md Ishtyaq Mahmud, Muntasir Mamun, and Ahmed Abdelgawad · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Fast submodular function maximization
Lianke Qin, Zhao Song, and Yitan Wang · 2023
Closest in time.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Closest in time.
A general algorithm for solving rank-one matrix sensing
Lianke Qin, Zhao Song, and Ruizhe Zhang · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Sketching for first order method: efficient algorithm for low-bandwidth channel and vulnerability
Zhao Song, Yitan Wang, Zheng Yu, and Lichen Zhang · 2023
Closest in time.
Sketching meets differential privacy: fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2023
Closest in time.
Efficient alternating minimization with applications to weighted low rank approximation
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Infoprompt: Information-theoretic soft prompt tuning for natural language understanding
Junda Wu, Tong Yu, Rui Wang, Zhao Song, Ruiyi Zhang, Handong Zhao, Chaochao Lu, Shuai Li, and Ricardo Henao · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
H _ 2 \_2 o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Closest in time.