Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Cited alongside, same era.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Original
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Original
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Original
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Going deeper with convolutions
Original
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Cited alongside, same era.