Fetching the paper…
Reading the bibliography…
The problem of efficient approximation of a linear operator induced by the Gaussian or softmax kernel is often addressed using random features (RFs) which yield an unbiased approximation of the operator's result.
Deep kernel learning via random Fourier features
Xie, J., Liu, F., Wang, K., and Huang, X · 1910
Earlier work this paper cites.
On estimating regression
Nadaraya, E. A · 1964
Earlier work this paper cites.
Smooth regression analysis
Watson, G. S · 1964
Earlier work this paper cites.
Dynamical systems that sort lists, diagonalize matrices, and solve linear programming problems
Brockett, R · 1991
Earlier work this paper cites.
Numerical Linear Algebra
Trefethen, L. N. and Bau, D · 1997
Earlier work this paper cites.
An elementary proof of a theorem of johnson and lindenstrauss
Dasgupta, S. and Gupta, A · 2003
Earlier work this paper cites.
Improved fast Gauss transform and efficient kernel density estimation
Yang, Duraiswami, Gumerov, and Davis · 2003
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2005
Earlier work this paper cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Davis, J., Sarlós, T., Belanger, D., Colwell, L. J., and Weller, A · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
Uniform approximation of functions with random bases
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Cho, Y. and Saul, L. K · 2009
Earlier work this paper cites.
Differentially private empirical risk minimization
Chaudhuri, K., Monteleoni, C., and Sarwate, A. D · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Nyström method vs random Fourier features: A theoretical and empirical comparison
Yang, T., Li, Y., Mahdavi, M., Jin, R., and Zhou, Z · 2012
Earlier work this paper cites.
An almost optimal unrestricted fast johnson-lindenstrauss transform
Ailon, N. and Liberty, E · 2013
Cited alongside, same era.
Fastfood - computing hilbert space expansions in loglinear time
Le, Q. V., Sarlós, T., and Smola, A. J · 2013
Cited alongside, same era.
Random Laplace feature maps for semigroup kernels on histograms
Yang, J., Sindhwani, V., Fan, Q., Avron, H., and Mahoney, M. W · 2014
Cited alongside, same era.
Large-scale random features for kernel regression
Laparra, V., Gonzalez, D. M., Tuia, D., and Camps-Valls, G · 2015
Cited alongside, same era.
Fast function to function regression
Oliva, J. B., Neiswanger, W., Póczos, B., Xing, E. P., Trac, H., Ho, S., and Schneider, J. G · 2015
Cited alongside, same era.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Cited alongside, same era.
Initialization matters: Orthogonal predictive state recurrent neural networks
Choromanski, K., Downey, C., and Boots, B · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
But how does it work in theory? Linear SVM with random features
Sun, Y., Gilbert, A. C., and Tewari, A · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spherical random features for polynomial kernels
Pennington, J., Yu, F. X., and Kumar, S · 2015
Cited alongside, same era.
Optimal rates for random Fourier features
Sriperumbudur, B. K. and Szabó, Z · 2015
Cited alongside, same era.
On the error of random Fourier features
Sutherland, D. J. and Schneider, J. G · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
Quasi-monte carlo feature maps for shift-invariant kernels
Avron, H., Sindhwani, V., Yang, J., and Mahoney, M. W · 2016
Cited alongside, same era.
Minh, H. Q · 2016
Cited alongside, same era.
Later among the works it cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Improved noisy student training for automatic speech recognition
Park, D. S., Zhang, Y., Jia, Y., Han, W., Chiu, C., Li, B., Wu, Y., and Le, Q. V · 2020
Later among the works it cites.
Nonparametric adaptive control and prediction: Theory and randomized algorithms
Boffi, N. M., Tu, S., and Slotine, J. E · 2021
Later among the works it cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Later among the works it cites.
Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality
Gonon, L · 2021
Later among the works it cites.
Random features for the neural tangent kernel
Han, I., Avron, H., Shoham, N., Kim, C., and Shin, J · 2021
Later among the works it cites.
Towards a unified analysis of random Fourier features
Li, Z., Ton, J., Oglic, D., and Sejdinovic, D · 2021
Later among the works it cites.
Hybrid random features
Choromanski, K., Chen, H., Lin, H., Ma, Y., Sehanobish, A., Jain, D., Ryoo, M. S., Varley, J., Zeng, A., Likhosherstov, V., Kalashnikov, D., Sindhwani, V., and Weller, A · 2022
Later among the works it cites.
On learning the transformer kernel
Chowdhury, S. P., Solomou, A., Dubey, A., and Sachan, M · 2022
Later among the works it cites.
Chefs’ random tables: Non-trigonometric random features
Likhosherstov, V., Choromanski, K., Dubey, A., Liu, F., Sarlos, T., and Weller, A · 2022
Later among the works it cites.