Fetching the paper…
Reading the bibliography…
TensorFlow is a machine learning system that operates at large scale and in heterogeneous environments.
Annual review of computer science vol. 1, 1986
Arvind and D. E. Culler · 1986
Earlier work this paper cites.
Learning distributed representations of concepts
G. E. Hinton · 1986
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
M. I. Jordan · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Torch: A modular machine learning software library
R. Collobert, S. Bengio, and J. Mariéthoz · 2002
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
J. Dean and S. Ghemawat · 2004
Earlier work this paper cites.
Web 1T 5-gram version 1, 2006
T. Brants and A. Franz · 2006
Earlier work this paper cites.
The Chubby lock service for loosely-coupled distributed systems
M. Burrows · 2006
Earlier work this paper cites.
Map-reduce for machine learning on multicore
C. tao Chu, S. K. Kim, Y. an Lin, Y. Yu, G. Bradski, K. Olukotun, and A. Y. Ng · 2007
Earlier work this paper cites.
DryadLINQ: A system for general-purpose distributed data-parallel computing using a high-level language
Y. Yu, M. Isard, D. Fetterly, M. Budiu, U. Erlingsson, P. K. Gunda, and J. Currey · 2008
Earlier work this paper cites.
Exploring strategies for training deep neural networks
H. Larochelle, Y. Bengio, J. Louradour, and P. Lamblin · 2009
Earlier work this paper cites.
ZooKeeper: Wait-free coordination for internet-scale systems
P. Hunt, M. Konar, F. P. Junqueira, and B. Reed · 2010
Earlier work this paper cites.
An architecture for parallel topic models
A. Smola and S. Narayanamurthy · 2010
Earlier work this paper cites.
Mesos: A platform for fine-grained resource sharing in the data center
B. Hindman, A. Konwinski, M. Zaharia, A. Ghodsi, A. D. Joseph, R. Katz, S. Shenker, and I. Stoica · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. E. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Q. Le, M. Ranzato, R. Monga, M. Devin, G. Corrado, K. Chen, J. Dean, and A. Ng · 2012
Earlier work this paper cites.
Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing
M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn · 2013
Earlier work this paper cites.
LINQits: Big data on little clients
E. S. Chung, J. D. Davis, and J. Lee · 2013
Earlier work this paper cites.
DeVISE: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Multilingual acoustic models using distributed deep neural networks
G. Heigold, V. Vanhoucke, A. Senior, P. Nguyen, M. Ranzato, M. Devin, and J. Dean · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Naiad: a timely dataflow system
D. G. Murray, F. McSherry, R. Isaacs, M. Isard, P. Barham, and M. Abadi · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. Amarasinghe · 2013
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
gemmlowp: a small self-contained low-precision GEMM library, 2015
B. Jacob et al · 2015
Later among the works it cites.
On using very large target vocabulary for neural machine translation
S. Jean, K. Cho, R. Memisevic, and Y. Bengio · 2015
Later among the works it cites.
Fast algorithms for convolutional neural networks
A. Lavin and S. Gray · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. J. Rossbach, Y. Yu, J. Currey, J.-P. Martin, and D. Fetterly · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton · 2013
Cited alongside, same era.
On rectified linear units for speech processing
M. D. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. Le, P. Nguyen, A. Senior, V. Vanhoucke, J. Dean, and G. E. Hinton · 2013
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2014
Cited alongside, same era.
cuDNN: Efficient primitives for deep learning
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Cited alongside, same era.
Project Adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Cited alongside, same era.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Scalability! But at what COST?
F. McSherry, M. Isard, and D. G. Murray · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al · 2015
Later among the works it cites.
Toward accelerating deep learning at scale using specialized logic
K. Ovtcharov, O. Ruwase, J.-Y. Kim, J. Fowers, K. Strauss, and E. Chung · 2015
Later among the works it cites.
Nvidia devtech blog post
K. Powell · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Later among the works it cites.
Large-scale cluster management at Google with Borg
A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. J. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Józefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mane, R. Monga, S. Moore, D. G. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. A. Tucker, V. Vanhoucke, V. Vasudevan, F. B. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2016
Closest in time.
Revisiting distributed synchronous SGD
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Closest in time.
convnet-benchmarks, 2016
S. Chintala · 2016
Closest in time.
GeePS: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing · 2016
Closest in time.
Google supercharges machine learning tasks with TPU custom chip, 2016
N. Jouppi · 2016
Closest in time.
Exploring the limits of language modeling
R. Józefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Closest in time.
SparkNet: Training deep networks in Spark
P. Moritz, R. Nishihara, I. Stoica, and M. I. Jordan · 2016
Closest in time.
Movidius announces Deep Learning Accelerator and Fathom software framework, 2016
Movidius Ltd · 2016
Closest in time.
neon, 2016
Nervana Systems · 2016
Closest in time.
NCCL: Optimized primitives for collective multi-gpu communication, 2016
NVIDIA Corporation · 2016
Closest in time.