Fetching the paper…
Reading the bibliography…
Many recent machine learning models rely on fine-grained dynamic control flow for training and inference.
A preliminary architecture for a basic data-flow processor
Dennis, J. B., and Misunas, D. P · 1975
Earlier work this paper cites.
Dataflow architectures
Arvind, and Culler, D. E · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1988
Earlier work this paper cites.
Executing a program on the MIT tagged-token dataflow architecture
Arvind, and Nikhil, R. S · 1990
Earlier work this paper cites.
The if-problem in automatic differentiation
Beck, T., and Fischer, H · 1994
Earlier work this paper cites.
Learning task-dependent distributed representations by backpropagation through structure
Goller, C., and Küchler, A · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., and Schmidhuber, J · 1997
Earlier work this paper cites.
Torch: A modular machine learning software library
Collobert, R., Bengio, S., and Mariéthoz, J · 2002
Earlier work this paper cites.
Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator
Pearlmutter, B. A., and Siskind, J. M · 2008
Earlier work this paper cites.
Theano: A CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y · 2010
Earlier work this paper cites.
CIEL: A universal execution engine for distributed data-flow computing
Murray, D. G., Schwarzkopf, M., Smowton, C., Smit, S., Madhavapeddy, A., and Hand, S · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., Warde-Farley, D., and Bengio, Y · 2012
Earlier work this paper cites.
The Tapenade automatic differentiation tool: principles, model, and specification
Hascoët, L., and Pascual, V · 2013
Earlier work this paper cites.
Naiad: a timely dataflow system
Murray, D. G., McSherry, F., Isaacs, R., Isard, M., Barham, P., and Abadi, M · 2013
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Cited alongside, same era.
Autograd: Reverse-mode differentiation of native python
Maclaurin, D., Duvenaud, D., and Adams, R. P · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Google supercharges machine learning tasks with TPU custom chip, 2016
Jouppi, N · 2016
Later among the works it cites.
Incremental, iterative data processing with timely dataflow
Murray, D. G., McSherry, F., Isard, M., Isaacs, R., Barham, P., and Abadi, M · 2016
Later among the works it cites.
CNTK: Microsoft’s open-source deep-learning toolkit
Seide, F., and Agarwal, A · 2016
Later among the works it cites.
Google’s Neural Machine Translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Later among the works it cites.
Documentation for imperative mode
Kudlur, M · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Zheng, X · 2016
Cited alongside, same era.
Globally normalized transition-based neural networks
Andor, D., Alberti, C., Weiss, D., Severyn, A., Presta, A., Ganchev, K., Petrov, S., and Collins, M · 2016
Cited alongside, same era.
Baydin, A. G., Pearlmutter, B. A., and Siskind, J. M · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Cited alongside, same era.
MXNet for deep learning, 2016
DMLC · 2016
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Cited alongside, same era.
Looks, M., Herreshoff, M., Hutchins, D., and Norvig, P · 2017
Later among the works it cites.
Real-time machine learning: The missing pieces
Nishihara, R., Moritz, P., Wang, S., Tumanov, A., Paul, W., Schleier-Smith, J., Liaw, R., Jordan, M. I., and Stoica, I · 2017
Later among the works it cites.
Eager execution: An imperative, define-by-run interface to TensorFlow, 2017
Shankar, A., and Dobson, W · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Later among the works it cites.
The matrix calculus you need for deep learning
Parr, T., and Howard, J · 2018
Closest in time.
Some principles of differentiable programming languages, 2018
Plotkin, G · 2018
Closest in time.
pytorch.org
PyTorch · 2018
Closest in time.
Chainer: A powerful, flexible and intuitive framework of neural networks, 2018
Tokui, S · 2018
Closest in time.