Fetching the paper…
Reading the bibliography…
We propose a new class of random feature methods for linearizing softmax and Gaussian kernels called hybrid random features (HRFs) that automatically adapt the quality of kernel estimation to provide most accurate approximation in the defined regions of interest.
Using the Nyström method to speed up kernel machines
Christopher K. I. Williams and Matthias W. Seeger · 2000
Earlier work this paper cites.
Approximation algorithms for MAX-3-CUT and other problems via complex semidefinite programming
Michel X. Goemans and David P. Williamson · 2003
Earlier work this paper cites.
Random features for kernel approximation: A survey in algorithms, theory, and beyond
Fanghui Liu, Xiaolin Huang, Yudong Chen, and Johan A. K. Suykens · 2004
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 2004
Earlier work this paper cites.
Kernel methods for measuring independence
Arthur Gretton, Ralf Herbrich, Alexander J. Smola, Olivier Bousquet, and Bernhard Schölkopf · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Analysis and extension of arc-cosine kernels for large margin classification
Y. Cho and L. K. Saul · 2012
Earlier work this paper cites.
Nyström method vs random Fourier features: A theoretical and empirical comparison
Tianbao Yang, Yu-Feng Li, Mehrdad Mahdavi, Rong Jin, and Zhi-Hua Zhou · 2012
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Less is more: Nyström computational regularization
Alessandro Rudi, Raffaello Camoriano, and Lorenzo Rosasco · 2015
Earlier work this paper cites.
Quasi-monte carlo feature maps for shift-invariant kernels
Haim Avron, Vikas Sindhwani, Jiyan Yang, and Michael W. Mahoney · 2016
Earlier work this paper cites.
Recycling randomness with structure for sublinear time kernel expansions
Krzysztof Choromanski and Vikas Sindhwani · 2016
Cited alongside, same era.
Orthogonal random features
Felix X. Yu, Ananda Theertha Suresh, Krzysztof Marcin Choromanski, Daniel N. Holtmann-Rice, and Sanjiv Kumar · 2016
Cited alongside, same era.
The unreasonable effectiveness of structured random orthogonal embeddings
Krzysztof Marcin Choromanski, Mark Rowland, and Adrian Weller · 2017
Cited alongside, same era.
Random features for compositional kernels
Amit Daniely, Roy Frostig, Vineet Gupta, and Yoram Singer · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Large-scale kernel methods for independence testing
Qinyi Zhang, Sarah Filippi, Arthur Gretton, and Dino Sejdinovic · 2018
Later among the works it cites.
Unifying orthogonal Monte Carlo methods
Krzysztof Choromanski, Mark Rowland, Wenyu Chen, and Adrian Weller · 2019
Later among the works it cites.
Sampled softmax with random Fourier features
Ankit Singh Rawat, Jiecao Chen, Felix X. Yu, Ananda Theertha Suresh, and Sanjiv Kumar · 2019
Later among the works it cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Using the Output Embedding to Improve Language Models, 2017
Ofir Press and Lior Wolf · 2017
Cited alongside, same era.
On Wasserstein two-sample testing and related families of nonparametric tests
Aaditya Ramdas, Nicolás García Trillos, and Marco Cuturi · 2017
Cited alongside, same era.
Adaptive sampled softmax with kernel based sampling
Guy Blanc and Steffen Rendle · 2018
Cited alongside, same era.
Geometrically coupled Monte Carlo sampling
Mark Rowland, Krzysztof Choromanski, François Chalus, Aldo Pacchiano, Tamás Sarlós, Richard E. Turner, and Adrian Weller · 2018
Cited alongside, same era.
URL http://www.unitree.cc/
Unitree Robotics
Cited in the paper.
Unlocking pixels for reinforcement learning via implicit attention
Krzysztof Choromanski, Deepali Jain, Jack Parker-Holder, Xingyou Song, Valerii Likhosherstov, Anirban Santara, Aldo Pacchiano, Yunhao Tang, and Adrian Weller
Cited in the paper.
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller · 2021
Closest in time.
Stable, fast and accurate: Kernelized attention with relative positional encoding, 2021
Shengjie Luo, Shanda Li, Tianle Cai, Di He, Dinglan Peng, Shuxin Zheng, Guolin Ke, Liwei Wang, and Tie-Yan Liu · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong · 2021
Closest in time.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Closest in time.