Fetching the paper…
Reading the bibliography…
We present a method to improve video description generation by modeling higher-order interactions between video frames and described concepts.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
Brox, T., Bruhn, A., Papenberg, N., and Weickert, J · 2004
Earlier work this paper cites.
Large displacement optical flow: Descriptor matching in variational motion estimation
Brox, T. and Malik, J · 2011
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, David L. and Dolan, William B · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, James and Bengio, Yoshua · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Guadarrama, Sergio, Krishnamoorthy, Niveda, Malkarnenkar, Girish, Venugopalan, Subhashini, Mooney, Raymond, Darrell, Trevor, and Saenko, Kate · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
Rohrbach, M., Qiu, W., Titov, I., Thater, S., Pinkal, M., and Schiele, B · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Graves, Alex and Jaitly, Navdeep · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Cited alongside, same era.
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Language models for image captioning: The quirks and what works
Devlin, Jacob, Cheng, Hao, Fang, Hao, Gupta, Saurabh, Deng, Li, He, Xiaodong, Zweig, Geoffrey, and Mitchell, Margaret · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
Yao, Li, Torabi, Atousa, Cho, Kyunghyun, Ballas, Nicolas, Pal, Christopher, Larochelle, Hugo, and Courville, Aaron · 2015
Later among the works it cites.
Delving deeper into convolutional networks for learning video representations
Ballas, Nicolas, Yao, Li, Pal, Chris, and Courville, Aaron C · 2016
Closest in time.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Closest in time.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, Justin, Karpathy, Andrej, and Fei-Fei, Li · 2016
Closest in time.
Reasoning about entailment with neural attention
Rocktäschel, Tim, Grefenstette, Edward, Hermann, Karl Moritz, Kociský, Tomás, and Blunsom, Phil · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long-term recurrent convolutional networks for visual recognition and description
Donahue, Jeff, Hendricks, Lisa Anne, Guadarrama, Sergio, Rohrbach, Marcus, Venugopalan, Subhashini, Saenko, Kate, and Darrell, Trevor · 2015
Cited alongside, same era.
From captions to visual concepts and back
Fang, Hao, Gupta, Saurabh, Iandola, Forrest, Srivastava, Rupesh K., Deng, Li, Dollar, Piotr, Gao, Jianfeng, He, Xiaodong, Mitchell, Margaret, Platt, John C., Lawrence Zitnick, C., and Zweig, Geoffrey · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2015
Cited alongside, same era.
Action recognition using visual attention
Sharma, Shikhar, Kiros, Ryan, and Salakhutdinov, Ruslan · 2015
Cited alongside, same era.
End-to-end memory networks
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
Tran, Du, Bourdev, Lubomir, Fergus, Rob, Torresani, Lorenzo, and Paluri, Manohar · 2015
Cited alongside, same era.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pan, Pingbo, Xu, Zhongwen, Yang, Yi, Wu, Fei, and Zhuang, Yueting
Cited in the paper.
Closest in time.
Beyond caption to narrative: Video captioning with multiple sentences
Shin, Andrew, Ohnishi, Katsunori, and Harada, Tatsuya · 2016
Closest in time.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, Gunnar A., Varol, Gül, Wang, Xiaolong, Farhadi, Ali, Laptev, Ivan, and Gupta, Abhinav · 2016
Closest in time.
Encode, review, and decode: Reviewer module for caption generation
Yang, Zhilin, Yuan, Ye, Wu, Yuexin, Salakhutdinov, Ruslan, and Cohen, William W · 2016
Closest in time.
Video paragraph captioning using hierarchical recurrent neural networks
Yu, Haonan, Wang, Jiang, Huang, Zhiheng, Yang, Yi, and Xu, Wei · 2016
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron, Salakhudinov, Ruslan, Zemel, Rich, and Bengio, Yoshua · 2057
Closest in time.