Fetching the paper…
Reading the bibliography…
Deep Learning (DL) algorithms are the central focus of modern machine learning systems.
Random sampling with a reservoir
Vitter, J. S · 1985
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality
Indyk, P. and Motwani, R · 1998
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Gionis, A., Indyk, P., and Motwani, R · 1999
Earlier work this paper cites.
Loss decomposition for fast learning in large output spaces
Yen, I. E.-H., Kale, S., Yu, F., Holtmann-Rice, D., Kumar, S., and Ravikumar, P · 1999
Earlier work this paper cites.
Very sparse random projections
Li, P., Hastie, T. J., and Church, K. W · 2006
Earlier work this paper cites.
Gpu computing
Owens, J. D., Houston, M., Luebke, D., Green, S., Stone, J. E., and Phillips, J. C · 2008
Earlier work this paper cites.
Using intel® vtune™ performance analyzer events/ratios & optimizing applications
Malladi, R. K · 2009
Earlier work this paper cites.
Avoiding cache thrashing due to private data placement in last-level cache for manycore scaling
Meng, J. and Skadron, K · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Transparent huge pages in 2.6.38
Corbet, J · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Detecting false sharing in openmp applications using the darwin framework
Wicaksono, B., Tolubaeva, M., and Chapman, B · 2011
Earlier work this paper cites.
The power of comparative reasoning
Yagnik, J., Strelow, D., Ross, D. A., and Lin, R.-s · 2011
Cited alongside, same era.
Adaptive dropout for training deep neural networks
Ba, J. and Frey, B · 2013
Cited alongside, same era.
Efficient virtual memory for big memory servers
Basu, A., Gandhi, J., Chang, J., Hill, M., and Swift, M · 2013
Cited alongside, same era.
Makhzani, A. and Frey, B · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Cited alongside, same era.
Performance analysis of the memory management unit under scale-out workloads
Karakostas, V., Unsal, O., Nemirovsky, M., Cristal, A., and Swift, M · 2014
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Later among the works it cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Later among the works it cites.
The new intel xeon scalable processor(formerly skylake-sp)
Kumar, A., Soltis, D., Esmer, I., Yoaz, I., and Kottapalli, S · 2017
Later among the works it cites.
Adaptive sampled softmax with kernel based sampling
Blanc, G. and Rendle, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Powers of tensors and fast matrix multiplication
Le Gall, F · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Jean, S., Cho, K., Memisevic, R., and Bengio, Y · 2015
Cited alongside, same era.
Winner-take-all autoencoders
Makhzani, A. and Frey, B. J · 2015
Cited alongside, same era.
Quick training of probabilistic neural nets by importance sampling
Bengio, Y. et al
Cited in the paper.
Densified winner take all (wta) hashing for sparse datasets
Chen, B. and Shrivastava, A · 2018
Later among the works it cites.
Unique entity estimation with application to the syrian conflict
Chen, B., Shrivastava, A., and Steorts, R. C · 2018
Later among the works it cites.
Auto-tuning tensorflow threading model for CPU backend
Hasabnis, N · 2018
Later among the works it cites.
Scaling-up split-merge mcmc with locality sensitive sampling (lss)
Luo, C. and Shrivastava, A · 2018
Later among the works it cites.
Randomized algorithms accelerated over cpu-gpu for ultra-high dimensional similarity search
Wang, Y., Shrivastava, A., Wang, J., and Ryu, J · 2018
Later among the works it cites.
Fast and accurate stochastic gradient estimation
Chen, B., Xu, Y., and Shrivastava, A · 2019
Closest in time.