Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
What value do explicit high level concepts have in vision to language problems?
Q. Wu, C. Shen, L. Liu, A. Dick, , and A. v. d. Hengel · 2016
Later among the works it cites.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Sca-cnn: spatial and channel-wise attention in convolutional networks for image captioning
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua · 2017
Later among the works it cites.
Very deep convolutional networks for text classification
A. Conneau, H. Schwenk, L. Barrault, and Y. Lecun · 2017
Later among the works it cites.
Semantic compositional networks for visual captioning
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Later among the works it cites.
A hierarchical approach for generating descriptive image paragraphs
J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei · 2017
Later among the works it cites.
Knowing when to look: adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Later among the works it cites.
Hierarchical multimodal lstm for dense visual-semantic embedding
Z. Niu, M. Zhou, L. Wang, X. Gao, and G. Hua · 2017
Later among the works it cites.
Skeleton key: Image captioning by skeleton-attribute decomposition
Y. Wang, Z. Lin, X. Shen, S. Cohen, and G. W. Cottrell · 2017
Later among the works it cites.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2017
Later among the works it cites.