Fetching the paper…
Reading the bibliography…
Dense video captioning aims to generate multiple associated captions with their temporal locations from the video.
“Fast temporal activity proposals for efficient detection of human actions in untrimmed videos,”
F. C. Heilbron, J. C. Niebles, and B. Ghanem, · 1923
Earlier work this paper cites.
“BLEU: a method for automatic evaluation of machine translation,”
K. Papineni, S. Roukos, T. Ward, and W. Zhu, · 2002
Earlier work this paper cites.
“METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,”
A. Lavie and A. Agarwal, · 2005
Earlier work this paper cites.
“Actions in context,”
M. Marszalek, I. Laptev, and C. Schmid, · 2009
Earlier work this paper cites.
“Translating video content to natural language descriptions,”
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele, · 2013
Earlier work this paper cites.
“Large-scale video classification with convolutional neural networks,”
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Learning spatiotemporal features with 3D convolutional networks,”
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, · 2015
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation,”
R. Vedantam, C. L. Zitnick, and D. Parikh, · 2015
Earlier work this paper cites.
“Translating videos to natural language using deep recurrent neural networks,”
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko, · 2015
Earlier work this paper cites.
“Sequence to sequence-video to text,”
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko, · 2015
Earlier work this paper cites.
“Describing videos by exploiting temporal structure,”
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, · 2015
Earlier work this paper cites.
“Jointly modeling embedding and translation to bridge video and language,”
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, · 2016
Earlier work this paper cites.
“Video paragraph captioning using hierarchical recurrent neural networks,”
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu, · 2016
Earlier work this paper cites.
“Daps: Deep action proposals for action understanding,”
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem, · 2016
Earlier work this paper cites.
“Temporal action localization in untrimmed videos via multi-stage cnns,”
Z. Shou, D. Wang, and S. Chang, · 2016
Earlier work this paper cites.
“SST: Single-stream temporal action proposals,”
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles, · 2017
Earlier work this paper cites.
“Video captioning with attention-based LSTM and semantic consistency,”
L. Gao, Z. Guo, H. Zhang, X. Xu, and H. T. Shen, · 2017
Earlier work this paper cites.
“Dense-captioning events in videos,”
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Video captioning with guidance of multimodal latent topics,”
S. Chen, J. Chen, Q. Jin, and A. Hauptmann, · 2017
Cited alongside, same era.
“Turn tap: Temporal unit regression network for temporal action proposals,”
J. Gao, Z. Yang, K. Chen, C. Sun, and R. Nevatia, · 2017
Cited alongside, same era.
“Temporal action detection with structured segment networks,”
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin, · 2017
Cited alongside, same era.
“Jointly localizing and describing events for dense video captioning,”
Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei, · 2018
Cited alongside, same era.
“Joint event detection and description in continuous video streams,”
H. Xu, B. Li, V. Ramanishka, L. Sigal, and K. Saenko, · 2019
Later among the works it cites.
Z. Zhang, D. Xu, W. Ouyang, and C. Tan, “Show, tell and summarize: Dense video captioning using visual cue aided sentence summarization,”
2019
Later among the works it cites.
“Adversarial inference for multi-sentence video description,”
J. S. Park, M. Rohrbach, T. Darrell, and A. Rohrbach, · 2019
Later among the works it cites.
“Video action transformer network,”
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, · 2019
Later among the works it cites.
“Bmn: Boundary-matching network for temporal action proposal generation,”
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen, · 2019
Later among the works it cites.
“Event-centric hierarchical representation for dense video captioning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Reconstruction network for video captioning,”
B. Wang, L. Ma, W. Zhang, and W. Liu, · 2018
Cited alongside, same era.
“Bidirectional attentive fusion with context gating for dense video captioning,”
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu, · 2018
Cited alongside, same era.
“Hierarchical context encoding for events captioning in videos,”
D. Yang and C. Yuan, · 2018
Cited alongside, same era.
“Towards automatic learning of procedures from web instructional videos,”
L. Zhou, C. Xu, and J. J. Corso, · 2018
Cited alongside, same era.
“End-to-end dense video captioning with masked transformer,”
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, · 2018
Cited alongside, same era.
“Bsn: Boundary sensitive network for temporal action proposal generation,”
T. Lin, X. Zhao, H. Su, C. Wang, and M. Yang · 2018
Cited alongside, same era.
T. Wang and H. Zheng and M. Yu and Q. Tian and H. Hu, · 2020
Later among the works it cites.
“Multi-modal dense video captioning,”
V. Iashin and E. Rahtu, · 2020
Later among the works it cites.
“A better use of audio-visual cues: Dense video captioning with bi-modal transformer,”
V. Iashin and E. Rahtu, · 2020
Later among the works it cites.
”An efficient framework for dense video captioning,”
M. Suin and A. N. Rajagopalan, · 2020
Later among the works it cites.
“End-to-end object detection with transformers,”
N. Carion, M. Francisco, S. Gabriel, U. Nicolas, K. Alexander and Z. Sergey, · 2020
Later among the works it cites.
“Soda: Story oriented dense video captioning evaluation framework,”
S. Fujita, T. Hirao, H. Kamigaito, M. Okumura, and M. Nagata, · 2020
Later among the works it cites.
“Fast learning of temporal action proposal via dense boundary generator,”
C. Lin, J. Li, Y. Wang, Y. Tai, D. Luo, Z. Cui, C. Wang, J. Li, F. Huang, and R. Ji, · 2020
Later among the works it cites.
“Learning texture transformer network for image super-resolution,”
F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo, · 2020
Later among the works it cites.
“Deformable detr: Deformable transformers for end-to-end object detection,”
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, · 2021
Closest in time.
“Pre-trained image processing transformer,”
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, · 2021
Closest in time.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, · 2021
Closest in time.
“Training data-efficient image transformers & distillation through attention,”
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, · 2021
Closest in time.