I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le, “Attention augmented convolutional networks,” CoRR , 2019. [Online]. Available: http://arxiv.org/abs/1904.09925
Original
1904
Earlier work this paper cites.
D. Mohapatra, V. K. Chippa, A. Raghunathan, and K. Roy, “Design of voltage-scalable meta-functions for approximate computing,” in Design, Automation Test in Europe , DATE, March 2011, pp. 1–6
2011
Earlier work this paper cites.
P. Ram and A. G. Gray, “Maximum inner-product search using cone trees,” in ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD, 2012
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in International Conference on Neural Information Processing Systems , NIPS, 2013
2013
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in International Conference on Architectural Support for Programming Languages and Operating Systems , ASPLOS, 2014
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” CoRR , 2014
2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka, “Neural turing machines,” CoRR , 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Conference on Empirical Methods in Natural Language Processing , EMNLP, 2014
2014
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu, “Recurrent models of visual attention,” in International Conference on Neural Information Processing Systems , NIPS, 2014
2014
Earlier work this paper cites.
A. Shrivastava and P. Li, “Asymmetric LSH (ALSH) for sublinear time maximum inner product search (MIPS),” in International Conference on Neural Information Processing Systems , NIPS, 2014
2014
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “Dadiannao: A machine-learning supercomputer,” in IEEE/ACM International Symposium on Microarchitecture , MICRO, 2014
2014
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International Conference on Machine Learning, ICML , 2015
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus et al. , “End-to-end memory networks,” in International Conference on Neural Information Processing Systems , NIPS, 2015
2015
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. Corrado, A. Davis, J. Dean, M. Devin et al. , “Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” CoRR , 2015
2015
Earlier work this paper cites.
Facebook, “The bAbI project,” https://research.fb.com/downloads/babi
2015
Earlier work this paper cites.
K. M. Hermann, T. Kociský, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom, “Teaching machines to read and comprehend,” in Advances in Neural Information Processing Systems , NIPS, 2015
2015
Earlier work this paper cites.
A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” in Conference on Empirical Methods in Natural Language Processing , EMNLP, 2015
2015
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spatial transformer networks,” in International Conference on Neural Information Processing Systems , NIPS, 2015
2015
Earlier work this paper cites.
A. Auvolat and P. Vincent, “Clustering is efficient for approximate maximum inner product search,” CoRR , 2015
2015
Earlier work this paper cites.
Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “Shidiannao: Shifting vision processing closer to the sensor,” in International Symposium on Computer Architecture , ISCA, 2015
2015
Earlier work this paper cites.
D. Liu, T. Chen, S. Liu, J. Zhou, S. Zhou, O. Teman, X. Feng, X. Zhou, and Y. Chen, “Pudiannao: A polyvalent machine learning accelerator,” in International Conference on Architectural Support for Programming Languages and Operating Systems , ASPLOS, 2015
2015
Earlier work this paper cites.
C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Optimizing FPGA-based accelerator design for deep convolutional neural networks,” in ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , FPGA, 2015
2015
Earlier work this paper cites.
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “MovieQA: Understanding Stories in Movies through Question-Answering,” in Conference on Computer Vision and Pattern Recognition , CVPR, 2016
2016
Earlier work this paper cites.
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwinska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou, A. P. Badia, K. M. Hermann, Y. Zwols, G. Ostrovski, A. Cain, H. King, C. Summerfield, P. Blunsom, K. Kavukcuoglu, and D. Hassabis, “Hybrid computing using a neural network with dynamic external memory,” Nature , 2016
2016
Earlier work this paper cites.
A. P. Parikh, O. Täckström, D. Das, and J. Uszkoreit, “A decomposable attention model for natural language inference,” in Empirical Methods in Natural Language Processing , EMNLP, 2016
2016
Earlier work this paper cites.