Fetching the paper…
Reading the bibliography…
The linear layer is one of the most pervasive modules in deep learning representations.
A fast cosine transform in one and two dimensions
Makhoul, John · 1980
Earlier work this paper cites.
Matrix Computations
Golub, Gene H. and Van Loan, Charles F · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Efficient parallel algorithms for optical computing with the DFT primitive
Reif, John and Tyagi, Akhilesh · 1997
Earlier work this paper cites.
Algorithmic design of diffractive optical systems for information processing
Müller-Quade, Jörn, Aagedal, Harald, Beth, Th, and Schmid, Michael · 1998
Earlier work this paper cites.
Decomposing a matrix into circulant and diagonal factors
Schmid, Michael, Steinwandt, Rainer, Müller-Quade, Jörn, Rötteler, Martin, and Beth, Thomas · 2000
Earlier work this paper cites.
Approximating ideal diffractive optical systems
Huhtanen, Marko · 2008
Earlier work this paper cites.
The Fast Johnson Lindenstrauss Transform and approximate nearest neighbors
Ailon, Nir and Chazelle, Bernard · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Predicting parameters in deep learning
Denil, Misha, Shakibi, Babak, Dinh, Laurent, Ranzato, Marc’Aurelio, and de Freitas, Nando · 2013
Earlier work this paper cites.
Fastfood – approximating kernel expansions in loglinear time
Le, Quoc, Sarlós, Tamás, and Smola, Alex · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Sainath, Tara N., Kingsbury, Brian, Sindhwani, Vikas, Arisoy, Ebru, and Ramabhadran, Bhuvana · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Xue, Jian, Li, Jinyu, and Gong, Yifan · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, Kyunghyun, Van Merriënboer, Bart, Gulcehre, Caglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Memory bounded deep convolutional networks
Collins, Maxwell D. and Kohli, Pushmeet · 2014
Cited alongside, same era.
Compressing deep convolutional networks using vector quantization
Gong, Yunchao, Liu, Liu, Yang, Ming, and Bourdev, Lubomir · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Jia, Yangqing, Shelhamer, Evan, Donahue, Jeff, Karayev, Sergey, Long, Jonathan, Girshick, Ross, Guadarrama, Sergio, and Darrell, Trevor · 2014
Towards trainable media: Using waves for neural network-style training
Hermans, Michiel and Vaerenbergh, Thomas Van · 2015
Closest in time.
Distilling the knowledge in a neural network
Hinton, Geoffrey E., Vinyals, Oriol, and Dean, Jeffrey · 2015
Closest in time.
Factoring matrices into the product of circulant and diagonal matrices
Huhtanen, Marko and Perämäki, Allan · 2015
Closest in time.
Sparse convolutional neural networks
Liu, Baoyuan, Wang, Min, Foroosh, Hassan, Tappen, Marshall, and Pensky, Marianna · 2015
Closest in time.
Tensorizing neural networks
Novikov, Alexander, Podoprikhin, Dmitry, Osokin, Anton, and Vetrov, Dmitry · 2015
Closest in time.
FitNets: Hints for thin deep nets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Speeding up neural networks for large scale classification using WTA hashing
Bakhtiary, Amir H., Lapedriza, Àgata, and Masip, David · 2015
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, Charles, Cornebise, Julien, Kavukcuoglu, Koray, and Wierstra, Daan · 2015
Cited alongside, same era.
Compressing neural networks with the hashing trick
Chen, Wenlin, Wilson, James T., Tyree, Stephen, Weinberger, Kilian Q., and Chen, Yixin · 2015
Cited alongside, same era.
An exploration of parameter redundancy in deep networks with circulant projections
Cheng, Yu, Yu, Felix X, Feris, R, Kumar, Sanjiv, Choudhary, Alok, and Chang, Shih-Fu · 2015
Cited alongside, same era.
Neural Turing machines
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2015
Cited alongside, same era.
Han, Song, Mao, Huizi, and Dally, William J
Cited in the paper.
Romero, Adriana, Ballas, Nicolas, Kahou, Samira Ebrahimi, Chassang, Antoine, Gatta, Carlo, and Bengio, Yoshua · 2015
Closest in time.
Random projections through multiple optical scattering: Approximating kernels at the speed of light
Saade, Alaa, Caltagirone, Francesco, Carron, Igor, Daudet, Laurent, Dremeau, Angelique, Gigan, Sylvain, and Krzakala, Florent · 2015
Closest in time.
Structured transforms for small-footprint deep learning
Sindhwani, Vikas, Sainath, Tara N, and Kumar, Sanjiv · 2015
Closest in time.
End-to-end memory networks
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron, Salakhutdinov, Ruslan, Zemel, Richard, and Bengio, Yoshua · 2015
Closest in time.
Deep fried convnets
Yang, Zichao, Moczulski, Marcin, Denil, Misha, de Freitas, Nando, Smola, Alex, Song, Le, and Wang, Ziyu · 2015
Closest in time.